• SIMD Considered Harmful

    From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 24 00:00:06 2026
    From Newsgroup: comp.arch


    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to::

    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET

    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
    {} in VEC is compiler telling HW that no registers inside the Loop are live outside of
    the Loop

    LOOP1 is the ADD-CMP-BC as 1 instruction
    LOOP1 can perform as many ADD-CMP-BC as the width of the Loop in a single cycle LOOP1 is a single cycle calculation

    While code remains "in" the LOOP FETCH-DECODE remains idle and Reservation Stations
    perform loop iterations.

    Now let us consider a Great-Big-Out-of-Order implementation of vVM in the execution
    pipeline::

    Assuming a 4-cycle LD cache hit and 4 cycle FMAC delay the execution of one iteration is::

    | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
    | LD R6 | cache | hit | align |
    | LD R7 | cache | hit | align |
    | ST AG | cache | hit | | ST R8 |
    | F M A C |

    or 9-cycles of latency. To get here the machine needs 3 AGEN units and at least {2 LD + 1 ST}
    or {3 LD-ST} units, 1 FMAC unit, and 1 LOOP unit.

    We now write this iteration as a single line::

    | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
    | Iteration |

    vVM can perform multiple iterations, starting several per cycle up to the number of
    calculation lanes in HW (limited by the small constant in VEC #small,{})::

    | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
    +8 | Iteration 1 |
    +8+1 | Iteration 2 |
    +8+2 | Iteration 3 |
    +8+3 | LD R6 | cache | hit | align Iteration 4 |
    +8+4 | LD R7 | cache | hit | align Iteration 5 |
    +8+5 | ST R8 | cache | Iteration 6 | ST R8 |
    +8+6 | Iteration 7 |
    +8+7 | Iteration 8 |
    +8 | Iteration 9 |
    +8+1 | Iteration 10 |
    +8+2 | Iteration 11 |
    +8+3 | LD R6 | cache | hit | align Iteration 12 |
    +8+4 | LD R7 | cache | hit | align Iteration 13 |
    +8+5 | ST R8 | cache | Iteration 14 | ST R8 |
    +8+6 | Iteration 15 |
    +8+7 | Iteration 16 |
    | ad-infinitum

    Where each LD and ST accesses a whole 64-Byte cache line, feeding 8 FMAC units with LD data,
    and consuming all 8 FMAC results as data to be stored. Due to non-aligned to cache line
    boundary issues, in general the AGEN part must run 1 access in front of where the buffers
    feed the FAMC units so cache line boundary crossings are penalized once. {One should observe
    that when LD R6 hits, ST R8 also hits (TLB too) because it is the same address, performed at
    the same time.} {vVM arranges that if the ST line must be replaced in cache, that the miss-
    buffers will remain capable of absorbing the results prior to writing to memory hierarchy.}

    vVM creates its own masking--so that when the last several iterations are not needed, they
    are not performed. So, there is no setup/maintenance of vector length <register>.

    In addition, but of more power-importance, is that the LOOP runs out of the reservation
    stations by adding a small index to the reservation station entries keyed to the loop
    iteration, keeping everything in data flow order on a per iteration basis. So, while the
    above Loop is running, FETCH-DECODE-INSERT is IDLE. In this sense, My 66000 GBOoO core
    only DECODEs 8 instructions compared to 163 for RISC-V and thousands for other ISAs. So,
    while calculation power is essentially identical, front-end power is greatly reduced.

    You could say LOOP1 is predicted to be taken, but it uses no predictor, and the LOOP
    Function Unit performs a comparison (per cycle) so the LOOP early outs at exactly the
    right time--so, FETCHed and DECODEd instructions are not thrown away--nor is a branch
    predictor state modified by the LOOP prediction {flow control within iterations will
    update branch predictor state as needed}

    And finally, vVM does this without any SW visible register file, making context switches
    fast, and Thread state in memory small. Nor does SW need to worry about the iteration count
    not being a multiple of the SIMD or Vector register length.

    vVM is fundamentally different than SIMD or Cray-like vectors because it vectorizes loops
    not instructions, But also that if an exception is raised, the person debugging the
    application sees a scalar calculation instead of a vector calculation--that is exceptions
    are "taken" with the IP pointing at the instruction which raised the exception, and the
    register file filled with data of "that iteration".

    One could postulate that there are 16 or even 32-lanes of calculations::
    a) x86 has shown power problems at 8-wide leading to core frequency reductions b) the memory units would have to access 2 or 4 cache lines per iteration
    c) the cache buffering would grow by 2|u to 4|u
    d) at some point one has to draw the line.

    Now as to Vectors versus SIMD versus vVM:
    a) adding vectors to a scalar ISA adds on the order of 300 instructions,
    b) adding SIMD to a scalar ISA adds on the order of 1000 instructions,
    c) adding vVM to a scalar ISA adds exactly 2 instructions.

    Can vVM do everything vectors or SIMD can--frankly no. On the other hand, vVM can perform
    mixed width calculations such as doubleword = word * byte + halfword--for which no SIMD
    ISA has found the OpCode space consumption viable. vVM can vectorize every leaf-function
    of str* and mem* C library calls. Given byte copy loop:

    char *strcpy(char *restrict dest, const char *restrict src)
    {
    char *ret = dest;
    while (*dest++ = *src++)
    ;
    return ret;
    }

    strcpy:
    ADD R2,R2,-R1 // make R1 index to R2
    VEC #0,{R1}
    LDUB R4,[R2,R1]
    STB R4,[R1]
    LOOP3 NE,R1,R4,#0 // LOOP type 3
    RET

    Given a 64-byte wide data-path (as in the above example 8 Double FP = 64-Bytes) each
    iteration of the loop will Load 64-bytes, Store 64-bytes. Considering this is a 4 (or 5)
    instruction loop in most scalar ISAs, the loop will be performed as if 256 instructions
    per cycle. Thus, one can write byte-by-byte algorithms and vVM performs them at fill data
    path width, generating its own masking, without a programmer revisiting the critical data
    movement functions each new chip iteration.

    The LOOPn subgroup can perform {counted loops, data terminated loops, and counted and data
    terminated loops}--strncmp is an example:

    int strncmp(const char* s1, const char* s2, size_t n)
    {
    while(n--)
    if(*s1!=*s2)
    return *s1 - *s2;
    else
    s1++, s2++;
    return 0;
    }

    strncmp:
    BNE0 R3,exit
    MOV R4,#0 // change n-- to j=0; j<n; j++
    VEC #0,{R4}
    LDUB R5,[R1,R4]
    LDUB R6,[R2,R4]
    LOOP3 NE,R4,R5,R6 // LOOP type 3
    ADD R1,R5,-R6
    RET
    exit: MOV R1,#0
    RET

    This 3 instruction loop can also run 64|u2 bytes per cycle--but since most strings being
    compared are rather short, this tends to run in 1 or at most 2 iterations in applications
    like symbol-table lookups.

    Summary:
    a) fewer total instructions
    b) fewer instructions in the loop
    c) mixed width calculations
    d) mixed width memory accesses
    e) no SIMD register File
    f) no Vector register File
    g) SW is unconcerned about Register width or depth
    h) only 2 instructions
    i) scalar debug model
    j) byte loops as fast as wide register loops
    k) ease of software maintenance

    Mitch

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Thu Sep 24 02:05:58 2026
    From Newsgroup: comp.arch

    On Thu, 24 Sep 2026 00:00:06 GMT, MitchAlsup wrote:

    The old paper
    https://www.sigarch.org/simd-instructions-considered-harmful/
    compares SIMD implementations with Vector implementations--as-if
    they were the only 2 games in town. ...

    and goes on to extoll the beneficial properties of Vectors over SIMD
    using RISC-V as exemplary. Let me pivot the conversation back to My
    66000 virtual-Vector-Method vVM (an alternative to both SIMD and
    Vectors).

    Would you agree, though, that SIMD is the least attractive of the
    three approaches?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Thu Sep 24 07:54:35 2026
    From Newsgroup: comp.arch

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to::

    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never loaded.


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1
    may be useful if you are doing a reduction and care about say overflow
    of intermediate values. Are there use use cases for values other than
    zero and one?
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 24 17:09:48 2026
    From Newsgroup: comp.arch


    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:

    On Thu, 24 Sep 2026 00:00:06 GMT, MitchAlsup wrote:

    The old paper
    https://www.sigarch.org/simd-instructions-considered-harmful/
    compares SIMD implementations with Vector implementations--as-if
    they were the only 2 games in town. ...

    and goes on to extoll the beneficial properties of Vectors over SIMD
    using RISC-V as exemplary. Let me pivot the conversation back to My
    66000 virtual-Vector-Method vVM (an alternative to both SIMD and
    Vectors).

    Would you agree, though, that SIMD is the least attractive of the
    three approaches?

    On instruction addition, yes; however SIMD is generally better at
    superluminal execution--1 SIMD instruction outside of any looping
    behavior.

    SIMD has arguably better cache characteristics, but for really
    large vectors, both are cache killers. My 66000 vectors have
    the ability to bypass the cache (LD/ST miss) to avoid strip-
    mining of cache data by use-once data.

    Cray-vectors are arguably better at instruction addition than SIMD
    and arguably inefficient when the element count of the vector
    is small (4).

    At one point in my architectural career, I was a fan of vectors (1980-1995-ish); at another I was a fan of SIMD (2000-2015),
    but both fell out of favor when I finally saw the damage they do
    to instruction sets. That was the genesis of vVM.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 24 19:48:23 2026
    From Newsgroup: comp.arch


    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to::

    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never loaded.
    Yes, typo; s/R2/R7/


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1
    may be useful if you are doing a reduction and care about say overflow
    of intermediate values. Are there use use cases for values other than
    zero and one?

    # not zero is how the compiler indicates a limit to the width of execution.
    So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3; so, the recurrence is easily processed.

    for( unsigned i=4; i<MAX; i++ )
    a[i] = a[i] + b[i]*a[i-3];

    This still vVM vectorizes but the width of iteration is limited to 3 so
    that successive loops have ready data to use as operands.
    {I am not saying this well (or in terminology of compiler writers)}

    So,
    #0 means as wide as HW has FUs
    #1 means 1-wide,
    #2 means 2-wide,
    #3 means 3-wide, ...




    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Thu Sep 24 15:04:19 2026
    From Newsgroup: comp.arch

    On 9/23/2026 9:05 PM, Lawrence DrCOOliveiro wrote:
    On Thu, 24 Sep 2026 00:00:06 GMT, MitchAlsup wrote:

    The old paper
    https://www.sigarch.org/simd-instructions-considered-harmful/
    compares SIMD implementations with Vector implementations--as-if
    they were the only 2 games in town. ...

    and goes on to extoll the beneficial properties of Vectors over SIMD
    using RISC-V as exemplary. Let me pivot the conversation back to My
    66000 virtual-Vector-Method vVM (an alternative to both SIMD and
    Vectors).

    Would you agree, though, that SIMD is the least attractive of the
    three approaches?

    SIMD is the cheapest and most effective option for short vectors (2 or 4 elements), but going much past this point, it scales poorly.

    If you want ever bigger SIMD vectors and to deal with lots of vector
    sizes and types, the traditional SIMD approach (new ops for every
    combination) becomes a terrible option.


    Meanwhile, something like a vector ISA makes more sense if one assumes
    that the data processing is primarily working on parallel arrays.

    It, however, sucks if you just sort of expect to use it to deal with
    random 2 and 4 element vectors spread here and there in the code.

    The latter could be reduced if rather than using a global vector state,
    there was sort of a collection of "vector presets" and the operation
    encodes which of the presets to use.

    One unavoidable cost with any sort of large vector ISA is that one will
    need to spend the cost of having a comparably large vector register file (well, either this, or stream vectors from memory; but then the design
    is limited by the memory ports, and potentially still has a high
    up-front cost).

    ...




    Otherwise...

    Implemented yet another experimental color-cell video format, and it
    seems to have been successful at Q/bpp and decode speed.

    General high level summary:
    Block formats:
    4x4x2, 4x4x1, 2x2x2, 2x2x1, 6-bit pattern table, flat color;
    Reuse selectors (from last 16 'new' blocks);
    Skip with translate (with an offset in 4x4 blocks).
    Color Endpoint Formats:
    RGB444Dy3
    2xRGB555 (most expensive option)
    Small Delta (5-bit index)
    Reused Endpoint (from last 16 endpoints)
    Dual Delta (Separate 5b delta for each endpoint)
    Dual Repeat (Selects low and high colors from last 16 endpoints)
    Uses 5b selectors, with 4b index, 1b to select color A or B.
    None (Block skips endpoint, reuses prior).


    Image data is encoded as a linear byte stream (no entropy coder).
    Does have an optional RP2 post compressor, may be used if the RP2
    compressor results in a 20% size reduction. Decided to skip bothering
    with LZ4 this time (the LZ4 results are almost invariably worse than RP2
    at this).

    A quick test (by zipping the AVIs) does imply that Deflate could gain
    around a 24% size reduction, but Deflate is generally too slow.


    New (vs last codec):
    The endpoint reuse and dual cases were recent additions.
    The dual cases are mostly used as a last-ditch effort when the
    alternative is falling back to a pair of RGB555 colors (worst case).

    Also new:
    The image is now broken up into 16x16 macroblocks, with the 4x4 sub
    blocks encoded in Hilbert order. This notably increases locality for
    endpoints and similar relative to plain raster, and also greatly reduces
    the obvious "block streaks" artifacts from many of my previous
    color-cell compressors (at lower quality levels).


    Currently getting OK results between around 0.2 to 0.5 bpp (though,
    depends a lot on the video). Mostly targeting fairly low resolutions ATM.

    Compression results seem to be beating out its predecessor (5B).

    Test videos like "Bad Apple" and "Heyyeayeayea" can be compressed down
    to single digit MB without looking too horrible (and roughly an order of magnitude smaller than the same videos as CRAM at the same resolution).


    Decode speed (on my Zen+):
    Generally between 700 to 1000 megapixels/second to RGB555
    Around 600 to RGBA32 ATM
    Majority of the time here is spent unpacking pixel blocks.

    Falls short of reaching the original target of CRAM like speeds, but
    does quite notably beat CRAM on Q / bpp.

    Encode speed is kinda slow ATM, but could be made faster. At present
    spends a lot of time mostly doing searches to figure out which block
    offsets are best for doing skips and similar.

    Maximum theoretical image quality is similar to DXT1.


    Note that on Q/bpp, it would still basically get beat hard by pretty
    much any of the MPEG-style codecs.

    But, am seemingly now at least a little closer (well, and without paying
    the steep speed cost of using Rice coding, as in some of my earlier
    designs).

    But, yeah.

    ...


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stephen Fuld@sfuld@alumni.cmu.edu.invalid to comp.arch on Thu Sep 24 13:58:54 2026
    From Newsgroup: comp.arch

    On 9/24/2026 12:48 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to::

    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never
    loaded.
    Yes, typo; s/R2/R7/


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1
    may be useful if you are doing a reduction and care about say overflow
    of intermediate values. Are there use use cases for values other than
    zero and one?

    # not zero is how the compiler indicates a limit to the width of execution. So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3; so, the recurrence is easily processed.

    for( unsigned i=4; i<MAX; i++ )
    a[i] = a[i] + b[i]*a[i-3];

    This still vVM vectorizes but the width of iteration is limited to 3 so
    that successive loops have ready data to use as operands.
    {I am not saying this well (or in terminology of compiler writers)}

    Well, I get it. That makes sense. But given the limited values that #
    can have, what happens if you have a loop dependency greater than the #
    field allows? i.e. the 3 in your example above is say 17?
    --
    - Stephen Fuld
    (e-mail address disguised to prevent spam)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 24 22:24:10 2026
    From Newsgroup: comp.arch


    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/24/2026 12:48 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to:: >>>
    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >> loaded.
    Yes, typo; s/R2/R7/


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1
    may be useful if you are doing a reduction and care about say overflow
    of intermediate values. Are there use use cases for values other than
    zero and one?

    # not zero is how the compiler indicates a limit to the width of execution. So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
    so, the recurrence is easily processed.

    for( unsigned i=4; i<MAX; i++ )
    a[i] = a[i] + b[i]*a[i-3];

    This still vVM vectorizes but the width of iteration is limited to 3 so that successive loops have ready data to use as operands.
    {I am not saying this well (or in terminology of compiler writers)}

    Well, I get it. That makes sense. But given the limited values that #
    can have, what happens if you have a loop dependency greater than the # field allows? i.e. the 3 in your example above is say 17?

    Well #0..#31 are available. -:)
    AND we are not really thinking about #k >= 8 (power more than area)

    So, first order::
    a) Cray-vector machines will not vectorize.
    b) SIMD-vector machines will not vectorize, either.
    c) Both run out of vRF at r[-17].
    d) so little expected damage to competitive perf.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Fri Sep 25 00:08:50 2026
    From Newsgroup: comp.arch

    On Thu, 24 Sep 2026 17:09:48 GMT, MitchAlsup wrote:

    Cray-vectors are arguably better at instruction addition than SIMD
    and arguably inefficient when the element count of the vector is
    small (4).

    In the Cray docs somewhere, it said the break-even point where it is
    better to go through the setup of a vector instruction rather then
    execute individual scalar instructions for each set of operands is a
    vector length of just 2.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Fri Sep 25 00:27:53 2026
    From Newsgroup: comp.arch


    MitchAlsup <user5857@newsgrouper.org.invalid> posted:


    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/24/2026 12:48 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to:: >>>
    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5<<3] // LOOP transfers control back to here >>> LDD R7,[R4,R5<<3]
    FMAC R8,R2,R6,R7 // fixed
    ST R8,[R4,R5<<3]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >> loaded.
    Yes, typo; s/R2/R7/

    Yes, typo, but R2 = a and it is to get multipllied by x[i] in R6

    I also got the memory reference indexing wrong !! ... grrrrrr


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1 >> may be useful if you are doing a reduction and care about say overflow >> of intermediate values. Are there use use cases for values other than >> zero and one?

    # not zero is how the compiler indicates a limit to the width of execution.
    So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
    so, the recurrence is easily processed.

    for( unsigned i=4; i<MAX; i++ )
    a[i] = a[i] + b[i]*a[i-3];

    This still vVM vectorizes but the width of iteration is limited to 3 so that successive loops have ready data to use as operands.
    {I am not saying this well (or in terminology of compiler writers)}

    Well, I get it. That makes sense. But given the limited values that # can have, what happens if you have a loop dependency greater than the # field allows? i.e. the 3 in your example above is say 17?

    Well #0..#31 are available. -:)
    AND we are not really thinking about #k >= 8 (power more than area)

    So, first order::
    a) Cray-vector machines will not vectorize.
    b) SIMD-vector machines will not vectorize, either.
    c) Both run out of vRF at r[-17].
    d) so little expected damage to competitive perf.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Fri Sep 25 00:41:53 2026
    From Newsgroup: comp.arch


    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:

    On Thu, 24 Sep 2026 17:09:48 GMT, MitchAlsup wrote:

    Cray-vectors are arguably better at instruction addition than SIMD
    and arguably inefficient when the element count of the vector is
    small (4).

    In the Cray docs somewhere, it said the break-even point where it is
    better to go through the setup of a vector instruction rather then
    execute individual scalar instructions for each set of operands is a
    vector length of just 2.

    Yes, when memory was 20-odd cycles away--but a 5 GHz machine will find
    memory 200 cycles away, not 20. That changes the break-even point (and
    the number of elements in a vector register).

    In addition, for vectors that short (4), caches work.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Fri Sep 25 17:44:59 2026
    From Newsgroup: comp.arch

    On 2026-Sep-24 18:24, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/24/2026 12:48 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to:: >>>>>
    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >>>> loaded.
    Yes, typo; s/R2/R7/


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1 >>>> may be useful if you are doing a reduction and care about say overflow >>>> of intermediate values. Are there use use cases for values other than >>>> zero and one?

    # not zero is how the compiler indicates a limit to the width of execution. >>> So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
    so, the recurrence is easily processed.

    for( unsigned i=4; i<MAX; i++ )
    a[i] = a[i] + b[i]*a[i-3];

    This still vVM vectorizes but the width of iteration is limited to 3 so
    that successive loops have ready data to use as operands.
    {I am not saying this well (or in terminology of compiler writers)}

    Well, I get it. That makes sense. But given the limited values that #
    can have, what happens if you have a loop dependency greater than the #
    field allows? i.e. the 3 in your example above is say 17?

    Well #0..#31 are available. -:)
    AND we are not really thinking about #k >= 8 (power more than area)

    So, first order::
    a) Cray-vector machines will not vectorize.
    b) SIMD-vector machines will not vectorize, either.
    c) Both run out of vRF at r[-17].
    d) so little expected damage to competitive perf.

    What does the vector width value do to the uArch
    (eg does it control the scheduler or a uOp packer)?
    Why would someone ever have a value other than #0 (autoconfig)?


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lawrence =?iso-8859-13?q?D=FFOliveiro?=@ldo@nz.invalid to comp.arch on Fri Sep 25 22:13:40 2026
    From Newsgroup: comp.arch

    On Fri, 25 Sep 2026 00:41:53 GMT, MitchAlsup wrote:

    On Fri, 25 Sep 2026 00:08:50 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    In the Cray docs somewhere, it said the break-even point where it
    is better to go through the setup of a vector instruction rather
    then execute individual scalar instructions for each set of
    operands is a vector length of just 2.

    Yes, when memory was 20-odd cycles away--but a 5 GHz machine will
    find memory 200 cycles away, not 20. That changes the break-even
    point (and the number of elements in a vector register).

    I donrCOt see why it does. The CPU-memory gap applies to the fetching of operands and storing of results back to memory, regardless of whether
    the same operation is being done by a scalar functional unit or a
    vector one. The extra overhead of setting up a vector operation
    happens entirely in the CPU, so should not be (significantly) affected
    by memory speeds.

    CPU speeds have become so high nowadays that it is often faster to
    compute a value again than it is to fetch the same previously-computed
    result from (cache) memory.

    In addition, for vectors that short (4), caches work.

    They would work just as well for the same number of operands to
    scalar operations, no better and no worse.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Fri Sep 25 23:03:39 2026
    From Newsgroup: comp.arch


    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-24 18:24, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/24/2026 12:48 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to:: >>>>>
    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >>>> loaded.
    Yes, typo; s/R2/R7/


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1 >>>> may be useful if you are doing a reduction and care about say overflow >>>> of intermediate values. Are there use use cases for values other than >>>> zero and one?

    # not zero is how the compiler indicates a limit to the width of execution.
    So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
    so, the recurrence is easily processed.

    for( unsigned i=4; i<MAX; i++ )
    a[i] = a[i] + b[i]*a[i-3];

    This still vVM vectorizes but the width of iteration is limited to 3 so >>> that successive loops have ready data to use as operands.
    {I am not saying this well (or in terminology of compiler writers)}

    Well, I get it. That makes sense. But given the limited values that #
    can have, what happens if you have a loop dependency greater than the #
    field allows? i.e. the 3 in your example above is say 17?

    Well #0..#31 are available. -:)
    AND we are not really thinking about #k >= 8 (power more than area)

    So, first order::
    a) Cray-vector machines will not vectorize.
    b) SIMD-vector machines will not vectorize, either.
    c) Both run out of vRF at r[-17].
    d) so little expected damage to competitive perf.

    What does the vector width value do to the uArch

    Nothing, it has to do with loop-recurrences with is architectural.

    (eg does it control the scheduler or a uOp packer)?
    Why would someone ever have a value other than #0 (autoconfig)?


    for( int i = 3; i < MAX; i++ )
    a[i] = x*b[i] + a[i-3];

    a[3] is dependent on a[0]
    a[4] is dependent of a[1]
    ...
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sat Sep 26 00:29:39 2026
    From Newsgroup: comp.arch


    Lawrence =?iso-8859-13?q?D=FFOliveiro?= <ldo@nz.invalid> posted:

    On Fri, 25 Sep 2026 00:41:53 GMT, MitchAlsup wrote:

    On Fri, 25 Sep 2026 00:08:50 -0000 (UTC), Lawrence DrCOOliveiro wrote:

    In the Cray docs somewhere, it said the break-even point where it
    is better to go through the setup of a vector instruction rather
    then execute individual scalar instructions for each set of
    operands is a vector length of just 2.

    Yes, when memory was 20-odd cycles away--but a 5 GHz machine will
    find memory 200 cycles away, not 20. That changes the break-even
    point (and the number of elements in a vector register).

    I donrCOt see why it does. The CPU-memory gap applies to the fetching of operands and storing of results back to memory, regardless of whether
    the same operation is being done by a scalar functional unit or a
    vector one. The extra overhead of setting up a vector operation
    happens entirely in the CPU, so should not be (significantly) affected
    by memory speeds.

    When memory was 20-cycles away, a 64-entry vector register could absorb
    the entire latency 3 times. If memory is 200-cycles away, a 64-entry
    vector register can absorb only 1/3 of a memory latency.

    Note we are talking about a core where we can send a new address to
    memory every clock, too.

    CPU speeds have become so high nowadays that it is often faster to
    compute a value again than it is to fetch the same previously-computed
    result from (cache) memory.

    In addition, for vectors that short (4), caches work.

    They would work just as well for the same number of operands to
    scalar operations, no better and no worse.

    Also note: the vectors as short as 2 are faster than scalars of length
    2 point (above) is for 20-cycle scalar LDs (no caches on CRAY.) So,
    LD-ADD-ST is 24 cycles, while VLD-VADD-VST is ~27 cycles {extra vector
    setup cycles}
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Sat Sep 26 06:41:01 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    When memory was 20-cycles away, a 64-entry vector register could absorb
    the entire latency 3 times. If memory is 200-cycles away, a 64-entry
    vector register can absorb only 1/3 of a memory latency.

    Modern microarchitectures have prefetchers and caches for absorbing
    DRAM latency.

    BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
    low latency. The best was around 50ns, but 70ns-100ns is more
    typical.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Sat Sep 26 11:26:11 2026
    From Newsgroup: comp.arch

    On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    e) no SIMD register File

    Unfortunately, this has two side effects:

    e1) More memory traffic
    e2) No wide permutes

    That means that a certain class of high-performance codes will be
    at a disadvantage against, let's say, an efficient AVX512
    code.

    Compilers are not able to reach this, but people who write
    assembler or use corresponding intrinsics can.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Terje Mathisen@terje.mathisen@tmsw.no to comp.arch on Sat Sep 26 17:57:03 2026
    From Newsgroup: comp.arch

    Thomas Koenig wrote:
    On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    e) no SIMD register File

    Unfortunately, this has two side effects:

    e1) More memory traffic
    e2) No wide permutes

    That means that a certain class of high-performance codes will be
    at a disadvantage against, let's say, an efficient AVX512
    code.

    Compilers are not able to reach this, but people who write
    assembler or use corresponding intrinsics can.

    Wide permutes aren't really needed when every visible operation is
    scalar, right?

    You might need a boatload more registers to be able to express exactly
    how 16 or 32 bytes should be permuted, and if the permutation is
    generated at runtime, then you're just out of luck afaik?

    Terje
    --
    - <Terje.Mathisen at tmsw.no>
    "almost all programming can be viewed as an exercise in caching"
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sat Sep 26 18:23:03 2026
    From Newsgroup: comp.arch


    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    When memory was 20-cycles away, a 64-entry vector register could absorb
    the entire latency 3 times. If memory is 200-cycles away, a 64-entry >vector register can absorb only 1/3 of a memory latency.

    Modern microarchitectures have prefetchers and caches for absorbing
    DRAM latency.

    BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
    low latency. The best was around 50ns, but 70ns-100ns is more
    typical.

    I wanted to use 300 cycles (60 ns); but felt people would think I
    was over playing my hand.


    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Sat Sep 26 13:52:02 2026
    From Newsgroup: comp.arch

    On 9/26/2026 6:26 AM, Thomas Koenig wrote:
    On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    e) no SIMD register File

    Unfortunately, this has two side effects:

    e1) More memory traffic
    e2) No wide permutes

    That means that a certain class of high-performance codes will be
    at a disadvantage against, let's say, an efficient AVX512
    code.


    Last I heard, AVX512 is still not widely supported.

    Though, looking, at least on the newest CPUs it seems to have moved into
    the realm of being fully supported.


    Meanwhile, the CPU I am still running is in the category of "Has AVX,
    but using AVX comes with a performance hit, because it is still 128-bit internally".

    There is always a risk that the newer CPUs that have AVX512 might fall
    into a similar category, say, breaking the vectors in half to use 256
    bit ops.


    Personally I remain skeptical:
    Even when AVX and AVX512 are supported and fast, there is likely only a
    subset of code that is meaningfully improved by them.

    And, this subset may be smaller than the subset where the compiler tries
    to vectorize things but then this actually makes the code slower.


    Compilers are not able to reach this, but people who write
    assembler or use corresponding intrinsics can.


    For the subset of code that can be mapped to or can make effective use
    of wide SIMD...

    Personally, I would rather see superscalar with narrower SIMD.
    Don't need 256-bit SIMD ops if one can run multiple 128-bit SIMD ops in parallel to the same basic effect (similar to what is typical on the
    integer paths).


    Though, this doesn't play as well with Intel style "well, we will make
    the SIMD registers wider" approach. Each time the registers get wider,
    then you need add more operations just to deal with it.


    Granted, the other option is "each new tier effectively halves the
    usable size of the register file", say:
    64x 64-bit
    32x 128-bit
    16x 256-bit
    8x 512-bit

    Though, if one doubles the size of the register file each time, this
    creates other hassles (say, needing a prefix to glue 2 or 3 bits onto
    each register field).

    But, then one can partly decouple "How big is my SIMD?" from "How big is
    my register file?".

    Well, and if your CPU somehow supports 256x or 512x 64-bit registers,
    well, spill and fill could be in premise eliminated. But, then one could almost need some new ABI wonk to deal with the issue that basically no functions actually need this many registers at the narrower size.

    Another option being that the effective size of the register space
    remains the same, but larger tiers may add more logical registers,
    meaning that if-used, they also need to use the larger operations to
    interact with them.


    Say:
    R0..R63: 64-bit access
    R64..R127: Not accessible to 64-bit ops, only 128-bit ops
    R128..R255: Not accessible to 128-bit ops, only 256-bit ops
    ...

    Implicitly, each extended half not needing to natively support access at narrower sizes unless the CPU adds some way to access them at the
    narrower size.

    Well, or maybe go the other way as well, defining a subset of registers
    where access is allowed at progressively narrower sizes.

    Say:
    R16..R31 may be split into 32x 32-bit, or 64x 16-bit
    R16..R23 may support 8-bit access as 64x 8-bit registers.

    Where, say, pretty much any such operation becomes a 3R1W Mask-MUX type
    of thing (implicitly selects the sub-element from the source registers,
    and performs an implicit masked-insert on the destination).


    This being not too far off from how one could support 64-bit operations
    on a processor with 128-bit register ports.

    Though, granted, not as extreme as what one would need to support direct 32/16/8-bit operations on partial registers (which is arguably more niche).

    Could maybe make sense for 32-bit though:
    'int' is still 32-bit, and still common;
    And, so allowing 'int' operations to use allocate half of a register
    could reduce the effective register pressure.

    Could add a penalty though every time one needs to use an operation
    outside the set of that supported by register halves.


    Though, could use a prefix case:
    Prefix both allows expressing the operation is on a register half, and provides additional selector bits. This avoids eating a bunch of
    encoding space, but does mean that using these half-registers
    effectively doubles the instruction size (but, then again, what could effectively be 128x 32-bit registers, or maybe 256x 16-bit...).

    Ironically, the same basic mechanism as what would be needed to support
    access to more registers if larger SIMD sizes made the register file bigger.

    If it glues 3 bits onto each register field (plus size selector), this supporting down to 8-bit access with 64 registers, or up to 512
    registers with 64-bit access.

    ...


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Sat Sep 26 14:47:01 2026
    From Newsgroup: comp.arch

    On 9/26/2026 1:23 PM, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    When memory was 20-cycles away, a 64-entry vector register could absorb
    the entire latency 3 times. If memory is 200-cycles away, a 64-entry
    vector register can absorb only 1/3 of a memory latency.

    Modern microarchitectures have prefetchers and caches for absorbing
    DRAM latency.

    BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
    low latency. The best was around 50ns, but 70ns-100ns is more
    typical.

    I wanted to use 300 cycles (60 ns); but felt people would think I
    was over playing my hand.


    Well, for me, it is like the weirdness of seeing videos where people
    claim read/write speeds on NVMe SSD's that seem to encroach on the
    territory I see when benchmarking "memcpy()" style operations.

    Well, and much faster than the 300 MB/s I see with a SATA SSD (where
    both the SSD and MOBO support SATA3, which theoretically has a higher
    limit, but in testing I see 300MB/s, or the SATA2 limit).


    But, for bulk RAM copy (within a single core), seemingly the CPUs cap
    out at around 3.4 GB/sec.

    Well, conversely without much difference here between memcpy and memset
    style tasks:
    Both memcpy and memset cap at around 3.4 GB/s.
    Even if memcpy also needs to read the memory.
    Almost as if external RAM loads and stores have different pipes.

    Then, seemingly, there is a system-level limit of somewhere around 24
    GB/s (with 4 RAM sticks at DDR4-2400 speeds on a Zen+ based PC).

    Limit is somewhat higher if operating within the realm of what fits in
    cache.


    Some of this differs from the behavior of my custom CPU core, where
    memset style operations are around twice the speed of memcpy.

    But, then again, would imagine the workings of the memory buses are
    likely quite different.

    ...


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sat Sep 26 21:24:30 2026
    From Newsgroup: comp.arch


    BGB <cr88192@gmail.com> posted:

    On 9/26/2026 6:26 AM, Thomas Koenig wrote:
    On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    e) no SIMD register File

    Unfortunately, this has two side effects:

    e1) More memory traffic
    e2) No wide permutes

    That means that a certain class of high-performance codes will be
    at a disadvantage against, let's say, an efficient AVX512
    code.

    --------------
    Though, this doesn't play as well with Intel style "well, we will make
    the SIMD registers wider" approach. Each time the registers get wider,
    then you need add more operations just to deal with it.


    Granted, the other option is "each new tier effectively halves the
    usable size of the register file", say:
    64x 64-bit
    32x 128-bit
    16x 256-bit
    8x 512-bit

    Though, if one doubles the size of the register file each time, this
    creates other hassles (say, needing a prefix to glue 2 or 3 bits onto
    each register field).

    The advantage of not having a SIMD file and/or not having a SW addressable SIMD-RF, is that each implementation can choose for itself the ideal width
    of the buffering (4..32) and appropriate depth of the buffering (64..512) without SW having to change its instructions.

    But, then one can partly decouple "How big is my SIMD?" from "How big is
    my register file?".

    My way you don't have to.

    Well, and if your CPU somehow supports 256x or 512x 64-bit registers,
    well, spill and fill could be in premise eliminated. But, then one could almost need some new ABI wonk to deal with the issue that basically no functions actually need this many registers at the narrower size.

    Another option being that the effective size of the register space
    remains the same, but larger tiers may add more logical registers,
    meaning that if-used, they also need to use the larger operations to interact with them.

    Consider the hassle of saving and restoring those at context switch time; versus not having to do anything at all. Code written for my model works
    on all implementations regardless of the width of the data-path. SW emits
    code to a vN model, HW figures out how wide it can run.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sat Sep 26 21:16:54 2026
    From Newsgroup: comp.arch


    Thomas Koenig <tkoenig@netcologne.de> posted:

    On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    e) no SIMD register File

    Unfortunately, this has two side effects:

    e1) More memory traffic

    unclear. Using VEC-LOOP, a LD instruction on a 6-wide machine will LD
    a whole cache line into a buffer. Other than not being SW addressable,
    the buffers provide permutations, and wide load without a visible RF.

    e2) No wide permutes

    Permutes are performed by LD/ST instructions (changing the order bytes
    are loaded or stored).

    That means that a certain class of high-performance codes will be
    at a disadvantage against, let's say, an efficient AVX512
    code.

    Yes.

    Compilers are not able to reach this, but people who write
    assembler or use corresponding intrinsics can.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sat Sep 26 21:31:07 2026
    From Newsgroup: comp.arch


    BGB <cr88192@gmail.com> posted:

    On 9/26/2026 1:23 PM, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    When memory was 20-cycles away, a 64-entry vector register could absorb >>> the entire latency 3 times. If memory is 200-cycles away, a 64-entry
    vector register can absorb only 1/3 of a memory latency.

    Modern microarchitectures have prefetchers and caches for absorbing
    DRAM latency.

    BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
    low latency. The best was around 50ns, but 70ns-100ns is more
    typical.

    I wanted to use 300 cycles (60 ns); but felt people would think I
    was over playing my hand.


    Well, for me, it is like the weirdness of seeing videos where people
    claim read/write speeds on NVMe SSD's that seem to encroach on the
    territory I see when benchmarking "memcpy()" style operations.

    The well published X-zillion 4KB random reads per second does not have
    any of those 4KB pages touched by CPU instructions ?? So, it measures
    only the PCIe perf and not memory perf.

    Well, and much faster than the 300 MB/s I see with a SATA SSD (where
    both the SSD and MOBO support SATA3, which theoretically has a higher
    limit, but in testing I see 300MB/s, or the SATA2 limit).

    It is kind of sad that a single link PCIe 5.0 can transfer all the data
    that 10-to-20 300MB SATA drives can produce/consume.

    But, for bulk RAM copy (within a single core), seemingly the CPUs cap
    out at around 3.4 GB/sec.

    One cache line every 4-8 cycles. Kind of sad, here, too. We should have
    on-Die interconnects of 2|u512-bits per cycle per step on the on-Die interconnect (Cache line, both ways, for each node).

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Sat Sep 26 20:56:50 2026
    From Newsgroup: comp.arch

    On 9/26/2026 4:31 PM, MitchAlsup wrote:

    BGB <cr88192@gmail.com> posted:

    On 9/26/2026 1:23 PM, MitchAlsup wrote:

    anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    When memory was 20-cycles away, a 64-entry vector register could absorb >>>>> the entire latency 3 times. If memory is 200-cycles away, a 64-entry >>>>> vector register can absorb only 1/3 of a memory latency.

    Modern microarchitectures have prefetchers and caches for absorbing
    DRAM latency.

    BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
    low latency. The best was around 50ns, but 70ns-100ns is more
    typical.

    I wanted to use 300 cycles (60 ns); but felt people would think I
    was over playing my hand.


    Well, for me, it is like the weirdness of seeing videos where people
    claim read/write speeds on NVMe SSD's that seem to encroach on the
    territory I see when benchmarking "memcpy()" style operations.

    The well published X-zillion 4KB random reads per second does not have
    any of those 4KB pages touched by CPU instructions ?? So, it measures
    only the PCIe perf and not memory perf.


    Yeah. I can't really help but feel at least a little skeptical about
    GB/s speeds being reported off of NVMe M.2 SSDs.

    Maybe if there is some special mechanism to move blocks between the SSD
    and RAM without needing to go through the CPU, or somehow multiple CPU
    cores working in parallel for SSD IO.


    OK, I guess the AI answer is partly that NVMe does at least partly
    bypass the CPU cores and goes more directly to/from RAM.


    Well, and much faster than the 300 MB/s I see with a SATA SSD (where
    both the SSD and MOBO support SATA3, which theoretically has a higher
    limit, but in testing I see 300MB/s, or the SATA2 limit).

    It is kind of sad that a single link PCIe 5.0 can transfer all the data
    that 10-to-20 300MB SATA drives can produce/consume.


    Still faster than HDDs, in any case.

    Also maybe an irony that one of the main use-case I am using LZ
    compression for: to accelerate IO reading operations from HDDs and
    SDcards, would be arguably useless with a modern SSD. Though, one would
    still have the use-case of getting more effective capacity out of
    storage devices.


    It is also ironic in a way:
    The traditional imagination is that of file/data compression being
    inherently slow; not something to be used as an effective force
    multiplier of the ~ 100 MB/s or so that one gets from reading data from
    an HDD or similar.

    But, then again, people in the compression space are more often
    obsessing on maximizing compression at the expense of speed, rather than "moderately effective but primarily speed oriented" designs (where a 1%
    space saving is only really justifiable if it can give a multi-percent
    boost in effective bandwidth).

    But, then again, in most cases one doesn't really want to save 1% off
    the size but pay a 10x slowdown in doing so.

    For reading data from an FS, one wants decompression to be fast. And,
    for reading/writing pagefile pages, both need to be fast.


    At present, not yet found an entropy coding scheme that can win at this
    game over "just use raw bytes".

    That said, something like static Huffman with a 12-bit symbol length
    limit can "almost" start to cross into "fast" territory.

    Say, something like:
    void ReadSymbolBlob(HuffContext *ctx, int hti, byte *dst, int len)
    {
    byte *cs, *ct;
    u16 *htab;
    u32 win;
    int pos, hte, l;

    cs=ctx->cs;
    pos=ctx->pos;
    htab=ctx->hufftab[hti];
    ct=dst;

    l=len;
    while(l--)
    {
    win=*(u32 *)cs;
    hte=htab[(win>>pos)&4095];
    pos+=hte>>12;
    *ct++=hte;
    cs+=(pos>>3);
    pos&=7;
    }
    ctx->cs=cs;
    ctx->pos=pos;
    }
    Then building the format around unpacking blobs of symbols in advance
    (and pulling bytes from the corresponding blob). This is awkward, but
    can help st least.

    ...


    But, for bulk RAM copy (within a single core), seemingly the CPUs cap
    out at around 3.4 GB/sec.

    One cache line every 4-8 cycles. Kind of sad, here, too. We should have on-Die interconnects of 2|u512-bits per cycle per step on the on-Die interconnect (Cache line, both ways, for each node).


    OK.


    I guess AI response is, that the particular combination of features I
    was seeing were more a thing specific to the Zen1/Zen+ thing, and
    wouldn't really generalize to Intel CPUs or to newer Zen CPUs (which
    have a higher cap on memcpy/memset bandwidth).


    But, OTOH, now is not a good time to consider PC upgrades...
    Prices are a bit crazy right now...


    Sort of reminds me of something else I saw in a video:
    Video was talking about EUV lithography, ASML, and all of the challenges
    that happened in trying to make it work, timelines, etc...

    Then the video finished out by saying something to the effect of "The
    device you are watching this on wouldn't exist without this tech...".

    Then looked, and in my case, both my PC and phone would be as-is,
    because *both* effectively use processes that predate EUV lithography.


    So, at most, it would be a world that hit a process wall of "basically
    the stuff I am already using...".


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Sun Sep 27 11:33:02 2026
    From Newsgroup: comp.arch

    On 2026-Sep-25 19:03, MitchAlsup wrote:

    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-24 18:24, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/24/2026 12:48 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to:: >>>>>>>
    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here >>>>>>> LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >>>>>> loaded.
    Yes, typo; s/R2/R7/


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1 >>>>>> may be useful if you are doing a reduction and care about say overflow >>>>>> of intermediate values. Are there use use cases for values other than >>>>>> zero and one?

    # not zero is how the compiler indicates a limit to the width of execution.
    So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
    so, the recurrence is easily processed.

    for( unsigned i=4; i<MAX; i++ )
    a[i] = a[i] + b[i]*a[i-3];

    This still vVM vectorizes but the width of iteration is limited to 3 so >>>>> that successive loops have ready data to use as operands.
    {I am not saying this well (or in terminology of compiler writers)}

    Well, I get it. That makes sense. But given the limited values that # >>>> can have, what happens if you have a loop dependency greater than the # >>>> field allows? i.e. the 3 in your example above is say 17?

    Well #0..#31 are available. -:)
    AND we are not really thinking about #k >= 8 (power more than area)

    So, first order::
    a) Cray-vector machines will not vectorize.
    b) SIMD-vector machines will not vectorize, either.
    c) Both run out of vRF at r[-17].
    d) so little expected damage to competitive perf.

    What does the vector width value do to the uArch

    Nothing, it has to do with loop-recurrences with is architectural.

    I'm confused as you said earlier "#0 in VEC is compiler telling HW
    that the loop can be executed as wide as HW has Lanes".

    (eg does it control the scheduler or a uOp packer)?
    Why would someone ever have a value other than #0 (autoconfig)?


    for( int i = 3; i < MAX; i++ )
    a[i] = x*b[i] + a[i-3];

    a[3] is dependent on a[0]
    a[4] is dependent of a[1]
    ...

    Any vector element overlap is handled by store-to-load forwarding
    (which is why VVM doesn't need the complex alias analysis
    of SIMD or its permute/shuffle lane editing instructions).

    I still don't see what the width# on the VEC instruction does
    to the hardware. In your example it should generate the same
    instructions except a LDD has an extra offset of -3*8

    for( int i = 3; i < MAX; i++ )
    a[i] = x*b[i] + a[i-3];

    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R2,[R4,R5,-24]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET

    What would change if the VEC #width was: #0, #2, #3, #4?


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 28 00:04:16 2026
    From Newsgroup: comp.arch


    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-25 19:03, MitchAlsup wrote:

    EricP <ThatWouldBeTelling@thevillage.com> posted:

    On 2026-Sep-24 18:24, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/24/2026 12:48 PM, MitchAlsup wrote:

    Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:

    On 9/23/2026 5:00 PM, MitchAlsup wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to::

    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here >>>>>>> LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET


    Typo? R7 is loaded but never used, and R2 is used in the FMAC but never
    loaded.
    Yes, typo; s/R2/R7/


    8 instructions total
    5 instructions in the Loop {LDD to LOOP1}

    Smaller than any of the ISAs mentioned in the paper.

    #0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes

    A question. I certainly see the advantages of #0, and I think that #1 >>>>>> may be useful if you are doing a reduction and care about say overflow >>>>>> of intermediate values. Are there use use cases for values other than >>>>>> zero and one?

    # not zero is how the compiler indicates a limit to the width of execution.
    So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
    so, the recurrence is easily processed.

    for( unsigned i=4; i<MAX; i++ )
    a[i] = a[i] + b[i]*a[i-3];

    This still vVM vectorizes but the width of iteration is limited to 3 so >>>>> that successive loops have ready data to use as operands.
    {I am not saying this well (or in terminology of compiler writers)} >>>>
    Well, I get it. That makes sense. But given the limited values that # >>>> can have, what happens if you have a loop dependency greater than the # >>>> field allows? i.e. the 3 in your example above is say 17?

    Well #0..#31 are available. -:)
    AND we are not really thinking about #k >= 8 (power more than area)

    So, first order::
    a) Cray-vector machines will not vectorize.
    b) SIMD-vector machines will not vectorize, either.
    c) Both run out of vRF at r[-17].
    d) so little expected damage to competitive perf.

    What does the vector width value do to the uArch

    Nothing, it has to do with loop-recurrences with is architectural.

    I'm confused as you said earlier "#0 in VEC is compiler telling HW
    that the loop can be executed as wide as HW has Lanes".

    (eg does it control the scheduler or a uOp packer)?
    Why would someone ever have a value other than #0 (autoconfig)?


    for( int i = 3; i < MAX; i++ )
    a[i] = x*b[i] + a[i-3];

    a[3] is dependent on a[0]
    a[4] is dependent of a[1]
    ...

    Any vector element overlap is handled by store-to-load forwarding
    (which is why VVM doesn't need the complex alias analysis
    of SIMD or its permute/shuffle lane editing instructions).

    One could write:

    for( int i = 3; i < MAX; i++ )
    {
    d3 = d2;
    d2 = d1;
    d1 = a[i] = x*b[i] + d3;
    }

    And change the ST->LD into R->R->R
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From aph@aph@littlepinkcloud.invalid to comp.arch on Mon Sep 28 09:53:36 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
    implementations with Vector implementations--as-if they were the only 2 games in town. The
    paper contains the <toy> benchmark::

    void daxpy(const size_t n, const double a, const double x[], double y[])
    {
    for (size_t i = 0; i < n; i++) {
    y[i] = a*x[i] + y[i];
    }
    }

    and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
    exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
    (an alternative to both SIMD and Vectors). The above code compiles to::

    daxpy:
    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R7,[R4,R5]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET

    Arm SVE gets you something similar, with a few assumptions such as x
    and y not overlapping:

    daxpy:
    mov w3, 0
    mov z0.d, d0
    whilelo p7.d, xzr, x0
    .L2:
    ld1d z31.d, p7/z, [x1, x3, lsl 3]
    ld1d z30.d, p7/z, [x2, x3, lsl 3]
    fmla z30.d, p7/m, z0.d, z31.d
    st1d z30.d, p7, [x2, x3, lsl 3]
    incd x3
    whilelo p7.d, x3, x0
    b.any .L2
    ret

    I guess this group's local terminology would classify that as Vector
    rather than SIMD, but I'm not sure I understand the distinction.

    Ansrew.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Mon Sep 28 15:29:46 2026
    From Newsgroup: comp.arch

    BGB <cr88192@gmail.com> writes:
    On 9/26/2026 4:31 PM, MitchAlsup wrote:


    The well published X-zillion 4KB random reads per second does not have
    any of those 4KB pages touched by CPU instructions ?? So, it measures
    only the PCIe perf and not memory perf.


    Yeah. I can't really help but feel at least a little skeptical about
    GB/s speeds being reported off of NVMe M.2 SSDs.

    Maybe if there is some special mechanism to move blocks between the SSD
    and RAM without needing to go through the CPU, or somehow multiple CPU
    cores working in parallel for SSD IO.

    Perhaps you're thinking of the six-plus decade old technique of
    Direct Memory Access (DMA) by peripherals?

    Given modern interconnect (cpu and device) technologies (e.g. PCIe) pretty
    much _all_ peripheral data is moved without CPU intervention.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Mon Sep 28 17:39:58 2026
    From Newsgroup: comp.arch

    On 9/28/2026 10:29 AM, Scott Lurndal wrote:
    BGB <cr88192@gmail.com> writes:
    On 9/26/2026 4:31 PM, MitchAlsup wrote:


    The well published X-zillion 4KB random reads per second does not have
    any of those 4KB pages touched by CPU instructions ?? So, it measures
    only the PCIe perf and not memory perf.


    Yeah. I can't really help but feel at least a little skeptical about
    GB/s speeds being reported off of NVMe M.2 SSDs.

    Maybe if there is some special mechanism to move blocks between the SSD
    and RAM without needing to go through the CPU, or somehow multiple CPU
    cores working in parallel for SSD IO.

    Perhaps you're thinking of the six-plus decade old technique of
    Direct Memory Access (DMA) by peripherals?

    Given modern interconnect (cpu and device) technologies (e.g. PCIe) pretty much _all_ peripheral data is moved without CPU intervention.


    That seems to be the sentiment to what I saw when I looked at it.

    Would have expected something like, say:
    SSD exposes IO buffers at some location in MMIO space or similar;
    CPU threads go in and copy the buffers, possibly driven by IRQ's that
    the data is ready.

    But, apparently it is something more like:
    OS tells the device to read some data and where to put it in RAM;
    PCIe device goes and does so, the PCIe bus hits the CPU but is deflected straight to the DDR controller.


    So, yeah, I guess in any case it is possible that these SSD's can move
    data faster than a single CPU core can do memcpy.

    ...



    Either way, quite different from how HDDs were once done on x86:
    Wait for IRQ;
    Do a bunch of IN/OUT instructions to move the data via IO ports.

    Well, and the trick of needing to double-pump the IDE requests to deal
    with HDDs having more than 128GB.


    Or, how SDcard works in my CPU core:
    Send data to be written via SPI to MMIO;
    Write a "do it" value into an MMIO control register;
    Pull until done;
    Read response from MMIO.

    Did end up eventually tweaking the SPI-MMIO interface to allow 32 bytes
    per request (QDATA0..QDATA3) because originally, the mechanism for
    pumping the bytes over SPI via MMIO had been the main bottleneck
    (originally, pumping bytes one at a time over MMIO capped out at around
    300 K/s; 32B per command capping out at around 10 MB/s; which is at
    least faster than the SPI link).


    Even with the patents now expired, ironically this didn't leave as much incentive to go for the "faster" modes that SDcards support (pushing 4
    data bits on each clock edge), because I would effectively need to
    design something like a more complicated special memory-mapped SDcard
    device to be able to make use of this (and also redesign the filesystem
    layer to be able to use a cache memory-mapped disk blocks).


    But, yeah...

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From scott@scott@slp53.sl.home (Scott Lurndal) to comp.arch on Tue Sep 29 00:09:15 2026
    From Newsgroup: comp.arch

    BGB <cr88192@gmail.com> writes:
    On 9/28/2026 10:29 AM, Scott Lurndal wrote:
    BGB <cr88192@gmail.com> writes:
    On 9/26/2026 4:31 PM, MitchAlsup wrote:


    Either way, quite different from how HDDs were once done on x86:
    Wait for IRQ;
    Do a bunch of IN/OUT instructions to move the data via IO ports.

    All the mainframes at that time had been using DMA for a couple
    of decades.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Tue Sep 29 12:17:02 2026
    From Newsgroup: comp.arch

    On 2026-Sep-27 20:04, MitchAlsup wrote:

    EricP <ThatWouldBeTelling@thevillage.com> posted:

    (eg does it control the scheduler or a uOp packer)?
    Why would someone ever have a value other than #0 (autoconfig)?


    for( int i = 3; i < MAX; i++ )
    a[i] = x*b[i] + a[i-3];

    a[3] is dependent on a[0]
    a[4] is dependent of a[1]
    ...

    Any vector element overlap is handled by store-to-load forwarding
    (which is why VVM doesn't need the complex alias analysis
    of SIMD or its permute/shuffle lane editing instructions).

    One could write:

    for( int i = 3; i < MAX; i++ )
    {
    d3 = d2;
    d2 = d1;
    d1 = a[i] = x*b[i] + d3;
    }

    And change the ST->LD into R->R->R

    I see the desirability of that but I don't see how HW can infer
    that reg-reg forwarding chain from the loop LD and ST instructions.
    Hmmm... perhaps using that the R4 base and R5 index are the same
    in the ST and LD then that indicates a forward.

    MOV R5,#0
    VEC #0,{}
    LDD R6,[R3,R5] // LOOP transfers control back to here
    LDD R2,[R4,R5,-24]
    FMAC R8,R2,R5,R6
    ST R8,[R4,R5]
    LOOP1 LT,R5,#1,R1
    RET

    It greatly complicates the physical register lifespan and freeing.
    I will use the notation Rn.v where n is the register number
    and .v is the loop version, so R5.3 is the 3rd version of R5.
    The FMAC creates R8.1, R8.2, R8.3. Normally the physical
    register assigned to R8.1 can be recovered when R8 is renamed
    again creating R8.2, and the instruction that renamed it retires,
    in this case when the FMAC R8.2 retires it can recover R8.1.
    However now it has to keep that physical register allocated
    even after FMAC R8.2 retires so it can be forwarded
    to LDD R2.3,[R4,R5,-24]. When LDD R2.3 retires it can
    recycle the physical register.

    Also if the loop is interrupted and restarted it would have
    to rebuild the forwarding chain and reinit the prior values.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Torbjorn Lindgren@tl@none.invalid to comp.arch on Tue Sep 29 17:01:04 2026
    From Newsgroup: comp.arch

    BGB <cr88192@gmail.com> wrote:
    Well, for me, it is like the weirdness of seeing videos where people
    claim read/write speeds on NVMe SSD's that seem to encroach on the
    territory I see when benchmarking "memcpy()" style operations.

    As others noted the CPU generally don't touch the data for that kind
    of test, it's basically a PCIe bandwidth & overhead test if the SSD is
    fast enough.

    One thing to remember is that for write the quoted speed won't be
    sustained for very long unless you're buying extremely expensive
    enterprise SSDs - most SSDs has a pseudo-SLC mode cache and then
    "folds" the data to TLC or QLC.

    Tom's Hardware always test SSDs "properly" (IMHO), so lets look at a
    review of random semi-recent (Dec 2025) high-end consumer SSD, the
    Corsair MP700 Pro XT [1].

    For the 2TB model Corsair lists up to 14900MB/s read and 14500MB/s
    write and someone mentioned it being capable of 3.3M IOPS but I can't
    find a quote for that.

    The ones Tom's lists is a bit lower at 12+GB/s and 2M IOPS but I have
    no doubt Corsair really did manage to get the numbers they advertise.

    But the section that I always also check is the "Sustained Write
    Performance and Cache Recovery".

    As usual for SSDs the sustained write performance is COMPLICATED - it
    pretty much always is unless the interface is very limiting (IE SATA).

    In this case the 2TB drive ran in pSLC mode at 13.7GB/s for 16 seconds
    (~220GB, using ~660GB of TLC flash cells), then switched to direct
    write to the TLC flash at 4.2GB/s but there's intermittent short
    period where the performance falls down to 1.96GB/s, showing the SSD
    has to do "folding" (convert the pSLC data to native TLC) to free
    flash storage.

    The pSLC mode is there to make most normal operations lightning fast,
    in most cases it'll have plenty of time to run the folding in the
    background so it's not noticeable.

    Most SSDs used dynamic pSLC cache using some (or all) of the free
    flash, so the max cache size will shrink if the SSD is full.

    I'm rather impressed by those numbers but then it's very much not a
    cheap drive.

    Once you get into cheaper drives the sustained write performance often
    tanks a LOT, glares at an old Samsung 870 QVO (SATA QLC from 2020)
    which when full has a sustained write speed of 35-40MB/s outside the
    now minute (full, remember) pSLC cache. And that's far from the worst
    QLC drive I've seen.

    Flash has gotten faster since then, when I checked a recent budget
    PCIe 5.0 x4 QLC drive the sustained write speed was shown as ~200MB/s
    which was about in my expected range.

    The manufacture advertises 5.0GB/s read and 4.2GB/s write speed for
    this drive and I have no doubt they're not lying about that, I'm sure
    it hits that advertised write speed until the pSLC cache runs full.

    And to be honest, for a lot of people it's probably go work fairly
    well at least if they don't fill it above say 80-90%. Yes, there will
    slowdowns when installing large programs and it's also a LOT cheaper
    than the premium TLC drive above.

    I do consider it misleading to not list the sustained write speed but
    I'm not listing the vendor because NO ONE list it except for high-end Enterprise SSDs.


    Well, and much faster than the 300 MB/s I see with a SATA SSD (where
    both the SSD and MOBO support SATA3, which theoretically has a higher
    limit, but in testing I see 300MB/s, or the SATA2 limit).

    I have plenty of older SATA MLC or TLC SSDs which will fill the
    interface in either read or write basically forever, because the
    native write speed to the flash significantly exceed the interface
    speed so there was no reason to implement a pSLC cache, it's all
    native writes.

    In practice this translates to 550-560 MB/s read or ~520 MB/write. If
    you're getting 300MB/s you're actually exceeding the capabilities of a
    SATA II link (even if only by a bit) :-)

    Not sure if anyone still manufacturer/sells "good" SATA SSDs?, the
    market has kind of moved on.

    1. https://www.tomshardware.com/pc-components/ssds/corsair-mp700-pro-xt-2tb-ssd-review
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Tue Sep 29 20:49:31 2026
    From Newsgroup: comp.arch

    On 9/29/2026 12:01 PM, Torbjorn Lindgren wrote:
    BGB <cr88192@gmail.com> wrote:
    Well, for me, it is like the weirdness of seeing videos where people
    claim read/write speeds on NVMe SSD's that seem to encroach on the
    territory I see when benchmarking "memcpy()" style operations.

    As others noted the CPU generally don't touch the data for that kind
    of test, it's basically a PCIe bandwidth & overhead test if the SSD is
    fast enough.

    One thing to remember is that for write the quoted speed won't be
    sustained for very long unless you're buying extremely expensive
    enterprise SSDs - most SSDs has a pseudo-SLC mode cache and then
    "folds" the data to TLC or QLC.

    Tom's Hardware always test SSDs "properly" (IMHO), so lets look at a
    review of random semi-recent (Dec 2025) high-end consumer SSD, the
    Corsair MP700 Pro XT [1].

    For the 2TB model Corsair lists up to 14900MB/s read and 14500MB/s
    write and someone mentioned it being capable of 3.3M IOPS but I can't
    find a quote for that.

    The ones Tom's lists is a bit lower at 12+GB/s and 2M IOPS but I have
    no doubt Corsair really did manage to get the numbers they advertise.

    But the section that I always also check is the "Sustained Write
    Performance and Cache Recovery".

    As usual for SSDs the sustained write performance is COMPLICATED - it
    pretty much always is unless the interface is very limiting (IE SATA).

    In this case the 2TB drive ran in pSLC mode at 13.7GB/s for 16 seconds (~220GB, using ~660GB of TLC flash cells), then switched to direct
    write to the TLC flash at 4.2GB/s but there's intermittent short
    period where the performance falls down to 1.96GB/s, showing the SSD
    has to do "folding" (convert the pSLC data to native TLC) to free
    flash storage.

    The pSLC mode is there to make most normal operations lightning fast,
    in most cases it'll have plenty of time to run the folding in the
    background so it's not noticeable.

    Most SSDs used dynamic pSLC cache using some (or all) of the free
    flash, so the max cache size will shrink if the SSD is full.

    I'm rather impressed by those numbers but then it's very much not a
    cheap drive.

    Once you get into cheaper drives the sustained write performance often
    tanks a LOT, glares at an old Samsung 870 QVO (SATA QLC from 2020)
    which when full has a sustained write speed of 35-40MB/s outside the
    now minute (full, remember) pSLC cache. And that's far from the worst
    QLC drive I've seen.

    Flash has gotten faster since then, when I checked a recent budget
    PCIe 5.0 x4 QLC drive the sustained write speed was shown as ~200MB/s
    which was about in my expected range.

    The manufacture advertises 5.0GB/s read and 4.2GB/s write speed for
    this drive and I have no doubt they're not lying about that, I'm sure
    it hits that advertised write speed until the pSLC cache runs full.

    And to be honest, for a lot of people it's probably go work fairly
    well at least if they don't fill it above say 80-90%. Yes, there will slowdowns when installing large programs and it's also a LOT cheaper
    than the premium TLC drive above.

    I do consider it misleading to not list the sustained write speed but
    I'm not listing the vendor because NO ONE list it except for high-end Enterprise SSDs.


    Yeah, I don't know, as I don't have one...


    But, either way, it implies IO speeds that somewhat exceed the
    single-core "memcpy()" on my PC.

    Then again, had noted not to long ago that the my main PC's
    single-threaded memcpy score is beaten by a laptop with a "Core i7 8th
    Gen".

    This laptop being fast at single-threaded tasks, though seems to more
    quickly lose speed with multiple threads.




    Well, and much faster than the 300 MB/s I see with a SATA SSD (where
    both the SSD and MOBO support SATA3, which theoretically has a higher
    limit, but in testing I see 300MB/s, or the SATA2 limit).

    I have plenty of older SATA MLC or TLC SSDs which will fill the
    interface in either read or write basically forever, because the
    native write speed to the flash significantly exceed the interface
    speed so there was no reason to implement a pSLC cache, it's all
    native writes.

    In practice this translates to 550-560 MB/s read or ~520 MB/write. If
    you're getting 300MB/s you're actually exceeding the capabilities of a
    SATA II link (even if only by a bit) :-)


    When I look at it in an HDD tool, it basically does a flat 300 MB/s with pretty much no variability.

    In this case, it was a fairly recent SSD made by Samsung to replace an
    older failing SSD (made by Sandisk). Had imaged over the contents from
    the old SSD to the new one.

    Drive is listed as "Samsung EVO 1TB".


    In this case, the HDD utility reports the drive interface as SATA 2.0
    (along with all the other SATA drives also being listed as SATA 2.0).

    Dunno the specifics, this is just what I seem to be seeing here.

    Theoretically, the MOBO chipset claims SATA 3.0 support.


    Checking:
    Power on hours: 18531, ~ 2 years.

    Curiously, read speed not really any faster than the old one, which also
    got 300 MB/s, but IIRC was slower at write speeds.


    Checking a few of my HDDs:
    6TB WD Red : 125-184 MB/s
    4TB WD Red : 169-178 MB/s
    2TB HGST Ultrastar: 180-201 MB/s


    The HGST drives being some second-hand (retired) server drives (still
    newer than the drives they replaced; *1). Currently, my 4TB WD Red is
    older than these drives, but my SSD and 6TB drive are younger.


    *1: IIRC around 116 and 100 kilohour.

    Though, does have an idle thought:
    Assuming a drive were hermetically sealed and used magnetic bearings and similar, how long could such a drive keep running?...

    Say, could one make a drive with megahour lifespans?...
    Say, a drive where humans expire faster than the drive wears out?...


    So, say:
    Drive is sealed, with an inert Ar atmosphere;
    Platter spindle is magnetically levitated during operation. Possibly the spindle uses a hybrid SRM design, with conical ends. These travel inside
    of aluminum conical indents, and when the rotor spins up it creates a repulsive force against the aluminum indents. Could maybe have a coating
    on the inside of the cones to deal with spin-up / spin-down; where the levitation would cease, but a fairly thin gap (to limit slop).
    Potentially, both surfaces are coated in tungsten or similar (for high abrasion resistence), and/or PTFE (where WC on WC would still be weak
    against sudden contact or vibration, and a WC/PTFE interface could
    absorb shock a little better). Well, or most of the interface is WC,
    with PTFE at the tips of the cones as this is the point where sudden
    contact is most likely (though a lot depends on acceleration vectors
    here). Well, or a WC+PTFE / W interface.

    Possibly, the magnet is moved to the arm, and the voice could is
    external, with the head arm powered inductively and using an optical
    link to the controller (to avoid metal fatigue in a flat-flex cable).
    Head-arm bushing could also be tungsten carbide or similar (on a
    tungsten shaft). Possibly the ends of the shaft (and head shaft bearing)
    use PFTE interfaces to reduce risk of damage from G-shock, ... Though,
    making the heads/platters resistant to G-shock being a harder problem.

    The LEDs/photodiodes would need to operate at a fairly low power though.

    Probably also use older and more robust process nodes so that metal
    migration and similar are less of an issue.


    But, alas, if built, would probably be too expensive...


    ...


    Not sure if anyone still manufacturer/sells "good" SATA SSDs?, the
    market has kind of moved on.

    1. https://www.tomshardware.com/pc-components/ssds/corsair-mp700-pro-xt-2tb-ssd-review

    Dunno there...


    --- Synchronet 3.22a-Linux NewsLink 1.2