The old paper
https://www.sigarch.org/simd-instructions-considered-harmful/
compares SIMD implementations with Vector implementations--as-if
they were the only 2 games in town. ...
and goes on to extoll the beneficial properties of Vectors over SIMD
using RISC-V as exemplary. Let me pivot the conversation back to My
66000 virtual-Vector-Method vVM (an alternative to both SIMD and
Vectors).
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to::
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here
LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
On Thu, 24 Sep 2026 00:00:06 GMT, MitchAlsup wrote:
The old paper
https://www.sigarch.org/simd-instructions-considered-harmful/
compares SIMD implementations with Vector implementations--as-if
they were the only 2 games in town. ...
and goes on to extoll the beneficial properties of Vectors over SIMD
using RISC-V as exemplary. Let me pivot the conversation back to My
66000 virtual-Vector-Method vVM (an alternative to both SIMD and
Vectors).
Would you agree, though, that SIMD is the least attractive of the
three approaches?
On 9/23/2026 5:00 PM, MitchAlsup wrote:
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to::
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here
LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
Typo? R7 is loaded but never used, and R2 is used in the FMAC but never loaded.Yes, typo; s/R2/R7/
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
A question. I certainly see the advantages of #0, and I think that #1
may be useful if you are doing a reduction and care about say overflow
of intermediate values. Are there use use cases for values other than
zero and one?
On Thu, 24 Sep 2026 00:00:06 GMT, MitchAlsup wrote:
The old paper
https://www.sigarch.org/simd-instructions-considered-harmful/
compares SIMD implementations with Vector implementations--as-if
they were the only 2 games in town. ...
and goes on to extoll the beneficial properties of Vectors over SIMD
using RISC-V as exemplary. Let me pivot the conversation back to My
66000 virtual-Vector-Method vVM (an alternative to both SIMD and
Vectors).
Would you agree, though, that SIMD is the least attractive of the
three approaches?
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 5:00 PM, MitchAlsup wrote:Yes, typo; s/R2/R7/
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to::
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here
LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
Typo? R7 is loaded but never used, and R2 is used in the FMAC but never
loaded.
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
A question. I certainly see the advantages of #0, and I think that #1
may be useful if you are doing a reduction and care about say overflow
of intermediate values. Are there use use cases for values other than
zero and one?
# not zero is how the compiler indicates a limit to the width of execution. So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3; so, the recurrence is easily processed.
for( unsigned i=4; i<MAX; i++ )
a[i] = a[i] + b[i]*a[i-3];
This still vVM vectorizes but the width of iteration is limited to 3 so
that successive loops have ready data to use as operands.
{I am not saying this well (or in terminology of compiler writers)}
On 9/24/2026 12:48 PM, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 5:00 PM, MitchAlsup wrote:Yes, typo; s/R2/R7/
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to:: >>>
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here
LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >> loaded.
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
A question. I certainly see the advantages of #0, and I think that #1
may be useful if you are doing a reduction and care about say overflow
of intermediate values. Are there use use cases for values other than
zero and one?
# not zero is how the compiler indicates a limit to the width of execution. So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
so, the recurrence is easily processed.
for( unsigned i=4; i<MAX; i++ )
a[i] = a[i] + b[i]*a[i-3];
This still vVM vectorizes but the width of iteration is limited to 3 so that successive loops have ready data to use as operands.
{I am not saying this well (or in terminology of compiler writers)}
Well, I get it. That makes sense. But given the limited values that #
can have, what happens if you have a loop dependency greater than the # field allows? i.e. the 3 in your example above is say 17?
Cray-vectors are arguably better at instruction addition than SIMD
and arguably inefficient when the element count of the vector is
small (4).
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/24/2026 12:48 PM, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 5:00 PM, MitchAlsup wrote:Yes, typo; s/R2/R7/
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to:: >>>
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5<<3] // LOOP transfers control back to here >>> LDD R7,[R4,R5<<3]
FMAC R8,R2,R6,R7 // fixed
ST R8,[R4,R5<<3]
LOOP1 LT,R5,#1,R1
RET
Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >> loaded.
--- Synchronet 3.22a-Linux NewsLink 1.2
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
A question. I certainly see the advantages of #0, and I think that #1 >> may be useful if you are doing a reduction and care about say overflow >> of intermediate values. Are there use use cases for values other than >> zero and one?
# not zero is how the compiler indicates a limit to the width of execution.
So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
so, the recurrence is easily processed.
for( unsigned i=4; i<MAX; i++ )
a[i] = a[i] + b[i]*a[i-3];
This still vVM vectorizes but the width of iteration is limited to 3 so that successive loops have ready data to use as operands.
{I am not saying this well (or in terminology of compiler writers)}
Well, I get it. That makes sense. But given the limited values that # can have, what happens if you have a loop dependency greater than the # field allows? i.e. the 3 in your example above is say 17?
Well #0..#31 are available. -:)
AND we are not really thinking about #k >= 8 (power more than area)
So, first order::
a) Cray-vector machines will not vectorize.
b) SIMD-vector machines will not vectorize, either.
c) Both run out of vRF at r[-17].
d) so little expected damage to competitive perf.
On Thu, 24 Sep 2026 17:09:48 GMT, MitchAlsup wrote:
Cray-vectors are arguably better at instruction addition than SIMD
and arguably inefficient when the element count of the vector is
small (4).
In the Cray docs somewhere, it said the break-even point where it is
better to go through the setup of a vector instruction rather then
execute individual scalar instructions for each set of operands is a
vector length of just 2.
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/24/2026 12:48 PM, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 5:00 PM, MitchAlsup wrote:Yes, typo; s/R2/R7/
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to:: >>>>>
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here
LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >>>> loaded.
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
A question. I certainly see the advantages of #0, and I think that #1 >>>> may be useful if you are doing a reduction and care about say overflow >>>> of intermediate values. Are there use use cases for values other than >>>> zero and one?
# not zero is how the compiler indicates a limit to the width of execution. >>> So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
so, the recurrence is easily processed.
for( unsigned i=4; i<MAX; i++ )
a[i] = a[i] + b[i]*a[i-3];
This still vVM vectorizes but the width of iteration is limited to 3 so
that successive loops have ready data to use as operands.
{I am not saying this well (or in terminology of compiler writers)}
Well, I get it. That makes sense. But given the limited values that #
can have, what happens if you have a loop dependency greater than the #
field allows? i.e. the 3 in your example above is say 17?
Well #0..#31 are available. -:)
AND we are not really thinking about #k >= 8 (power more than area)
So, first order::
a) Cray-vector machines will not vectorize.
b) SIMD-vector machines will not vectorize, either.
c) Both run out of vRF at r[-17].
d) so little expected damage to competitive perf.
On Fri, 25 Sep 2026 00:08:50 -0000 (UTC), Lawrence DrCOOliveiro wrote:
In the Cray docs somewhere, it said the break-even point where it
is better to go through the setup of a vector instruction rather
then execute individual scalar instructions for each set of
operands is a vector length of just 2.
Yes, when memory was 20-odd cycles away--but a 5 GHz machine will
find memory 200 cycles away, not 20. That changes the break-even
point (and the number of elements in a vector register).
In addition, for vectors that short (4), caches work.
On 2026-Sep-24 18:24, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/24/2026 12:48 PM, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 5:00 PM, MitchAlsup wrote:Yes, typo; s/R2/R7/
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to:: >>>>>
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here
LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >>>> loaded.
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
A question. I certainly see the advantages of #0, and I think that #1 >>>> may be useful if you are doing a reduction and care about say overflow >>>> of intermediate values. Are there use use cases for values other than >>>> zero and one?
# not zero is how the compiler indicates a limit to the width of execution.
So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
so, the recurrence is easily processed.
for( unsigned i=4; i<MAX; i++ )
a[i] = a[i] + b[i]*a[i-3];
This still vVM vectorizes but the width of iteration is limited to 3 so >>> that successive loops have ready data to use as operands.
{I am not saying this well (or in terminology of compiler writers)}
Well, I get it. That makes sense. But given the limited values that #
can have, what happens if you have a loop dependency greater than the #
field allows? i.e. the 3 in your example above is say 17?
Well #0..#31 are available. -:)
AND we are not really thinking about #k >= 8 (power more than area)
So, first order::
a) Cray-vector machines will not vectorize.
b) SIMD-vector machines will not vectorize, either.
c) Both run out of vRF at r[-17].
d) so little expected damage to competitive perf.
What does the vector width value do to the uArch
(eg does it control the scheduler or a uOp packer)?
Why would someone ever have a value other than #0 (autoconfig)?
On Fri, 25 Sep 2026 00:41:53 GMT, MitchAlsup wrote:
On Fri, 25 Sep 2026 00:08:50 -0000 (UTC), Lawrence DrCOOliveiro wrote:
In the Cray docs somewhere, it said the break-even point where it
is better to go through the setup of a vector instruction rather
then execute individual scalar instructions for each set of
operands is a vector length of just 2.
Yes, when memory was 20-odd cycles away--but a 5 GHz machine will
find memory 200 cycles away, not 20. That changes the break-even
point (and the number of elements in a vector register).
I donrCOt see why it does. The CPU-memory gap applies to the fetching of operands and storing of results back to memory, regardless of whether
the same operation is being done by a scalar functional unit or a
vector one. The extra overhead of setting up a vector operation
happens entirely in the CPU, so should not be (significantly) affected
by memory speeds.
CPU speeds have become so high nowadays that it is often faster to
compute a value again than it is to fetch the same previously-computed
result from (cache) memory.
In addition, for vectors that short (4), caches work.
They would work just as well for the same number of operands to
scalar operations, no better and no worse.
When memory was 20-cycles away, a 64-entry vector register could absorb
the entire latency 3 times. If memory is 200-cycles away, a 64-entry
vector register can absorb only 1/3 of a memory latency.
e) no SIMD register File
On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
e) no SIMD register File
Unfortunately, this has two side effects:
e1) More memory traffic
e2) No wide permutes
That means that a certain class of high-performance codes will be
at a disadvantage against, let's say, an efficient AVX512
code.
Compilers are not able to reach this, but people who write
assembler or use corresponding intrinsics can.
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
When memory was 20-cycles away, a 64-entry vector register could absorb
the entire latency 3 times. If memory is 200-cycles away, a 64-entry >vector register can absorb only 1/3 of a memory latency.
Modern microarchitectures have prefetchers and caches for absorbing
DRAM latency.
BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
low latency. The best was around 50ns, but 70ns-100ns is more
typical.
- anton--- Synchronet 3.22a-Linux NewsLink 1.2
On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
e) no SIMD register File
Unfortunately, this has two side effects:
e1) More memory traffic
e2) No wide permutes
That means that a certain class of high-performance codes will be
at a disadvantage against, let's say, an efficient AVX512
code.
Compilers are not able to reach this, but people who write
assembler or use corresponding intrinsics can.
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
When memory was 20-cycles away, a 64-entry vector register could absorb
the entire latency 3 times. If memory is 200-cycles away, a 64-entry
vector register can absorb only 1/3 of a memory latency.
Modern microarchitectures have prefetchers and caches for absorbing
DRAM latency.
BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
low latency. The best was around 50ns, but 70ns-100ns is more
typical.
I wanted to use 300 cycles (60 ns); but felt people would think I
was over playing my hand.
On 9/26/2026 6:26 AM, Thomas Koenig wrote:
On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
e) no SIMD register File
Unfortunately, this has two side effects:
e1) More memory traffic
e2) No wide permutes
That means that a certain class of high-performance codes will be
at a disadvantage against, let's say, an efficient AVX512
code.
--------------
Though, this doesn't play as well with Intel style "well, we will make
the SIMD registers wider" approach. Each time the registers get wider,
then you need add more operations just to deal with it.
Granted, the other option is "each new tier effectively halves the
usable size of the register file", say:
64x 64-bit
32x 128-bit
16x 256-bit
8x 512-bit
Though, if one doubles the size of the register file each time, this
creates other hassles (say, needing a prefix to glue 2 or 3 bits onto
each register field).
But, then one can partly decouple "How big is my SIMD?" from "How big is
my register file?".
Well, and if your CPU somehow supports 256x or 512x 64-bit registers,
well, spill and fill could be in premise eliminated. But, then one could almost need some new ABI wonk to deal with the issue that basically no functions actually need this many registers at the narrower size.
Another option being that the effective size of the register space
remains the same, but larger tiers may add more logical registers,
meaning that if-used, they also need to use the larger operations to interact with them.
On 2026-09-24, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
e) no SIMD register File
Unfortunately, this has two side effects:
e1) More memory traffic
e2) No wide permutes
That means that a certain class of high-performance codes will be
at a disadvantage against, let's say, an efficient AVX512
code.
Compilers are not able to reach this, but people who write
assembler or use corresponding intrinsics can.
On 9/26/2026 1:23 PM, MitchAlsup wrote:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
When memory was 20-cycles away, a 64-entry vector register could absorb >>> the entire latency 3 times. If memory is 200-cycles away, a 64-entry
vector register can absorb only 1/3 of a memory latency.
Modern microarchitectures have prefetchers and caches for absorbing
DRAM latency.
BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
low latency. The best was around 50ns, but 70ns-100ns is more
typical.
I wanted to use 300 cycles (60 ns); but felt people would think I
was over playing my hand.
Well, for me, it is like the weirdness of seeing videos where people
claim read/write speeds on NVMe SSD's that seem to encroach on the
territory I see when benchmarking "memcpy()" style operations.
Well, and much faster than the 300 MB/s I see with a SATA SSD (where
both the SSD and MOBO support SATA3, which theoretically has a higher
limit, but in testing I see 300MB/s, or the SATA2 limit).
But, for bulk RAM copy (within a single core), seemingly the CPUs cap
out at around 3.4 GB/sec.
BGB <cr88192@gmail.com> posted:
On 9/26/2026 1:23 PM, MitchAlsup wrote:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
MitchAlsup <user5857@newsgrouper.org.invalid> writes:
When memory was 20-cycles away, a 64-entry vector register could absorb >>>>> the entire latency 3 times. If memory is 200-cycles away, a 64-entry >>>>> vector register can absorb only 1/3 of a memory latency.
Modern microarchitectures have prefetchers and caches for absorbing
DRAM latency.
BTW, 200 cycles at 5 GHz would mean 40ns. I have never seen such a
low latency. The best was around 50ns, but 70ns-100ns is more
typical.
I wanted to use 300 cycles (60 ns); but felt people would think I
was over playing my hand.
Well, for me, it is like the weirdness of seeing videos where people
claim read/write speeds on NVMe SSD's that seem to encroach on the
territory I see when benchmarking "memcpy()" style operations.
The well published X-zillion 4KB random reads per second does not have
any of those 4KB pages touched by CPU instructions ?? So, it measures
only the PCIe perf and not memory perf.
Well, and much faster than the 300 MB/s I see with a SATA SSD (where
both the SSD and MOBO support SATA3, which theoretically has a higher
limit, but in testing I see 300MB/s, or the SATA2 limit).
It is kind of sad that a single link PCIe 5.0 can transfer all the data
that 10-to-20 300MB SATA drives can produce/consume.
But, for bulk RAM copy (within a single core), seemingly the CPUs cap
out at around 3.4 GB/sec.
One cache line every 4-8 cycles. Kind of sad, here, too. We should have on-Die interconnects of 2|u512-bits per cycle per step on the on-Die interconnect (Cache line, both ways, for each node).
EricP <ThatWouldBeTelling@thevillage.com> posted:
On 2026-Sep-24 18:24, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/24/2026 12:48 PM, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 5:00 PM, MitchAlsup wrote:Yes, typo; s/R2/R7/
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to:: >>>>>>>
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here >>>>>>> LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
Typo? R7 is loaded but never used, and R2 is used in the FMAC but never >>>>>> loaded.
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
A question. I certainly see the advantages of #0, and I think that #1 >>>>>> may be useful if you are doing a reduction and care about say overflow >>>>>> of intermediate values. Are there use use cases for values other than >>>>>> zero and one?
# not zero is how the compiler indicates a limit to the width of execution.
So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
so, the recurrence is easily processed.
for( unsigned i=4; i<MAX; i++ )
a[i] = a[i] + b[i]*a[i-3];
This still vVM vectorizes but the width of iteration is limited to 3 so >>>>> that successive loops have ready data to use as operands.
{I am not saying this well (or in terminology of compiler writers)}
Well, I get it. That makes sense. But given the limited values that # >>>> can have, what happens if you have a loop dependency greater than the # >>>> field allows? i.e. the 3 in your example above is say 17?
Well #0..#31 are available. -:)
AND we are not really thinking about #k >= 8 (power more than area)
So, first order::
a) Cray-vector machines will not vectorize.
b) SIMD-vector machines will not vectorize, either.
c) Both run out of vRF at r[-17].
d) so little expected damage to competitive perf.
What does the vector width value do to the uArch
Nothing, it has to do with loop-recurrences with is architectural.
(eg does it control the scheduler or a uOp packer)?
Why would someone ever have a value other than #0 (autoconfig)?
for( int i = 3; i < MAX; i++ )
a[i] = x*b[i] + a[i-3];
a[3] is dependent on a[0]
a[4] is dependent of a[1]
...
On 2026-Sep-25 19:03, MitchAlsup wrote:
EricP <ThatWouldBeTelling@thevillage.com> posted:
On 2026-Sep-24 18:24, MitchAlsup wrote:
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/24/2026 12:48 PM, MitchAlsup wrote:
Well, I get it. That makes sense. But given the limited values that # >>>> can have, what happens if you have a loop dependency greater than the # >>>> field allows? i.e. the 3 in your example above is say 17?
Stephen Fuld <sfuld@alumni.cmu.edu.invalid> posted:
On 9/23/2026 5:00 PM, MitchAlsup wrote:Yes, typo; s/R2/R7/
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to::
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here >>>>>>> LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
Typo? R7 is loaded but never used, and R2 is used in the FMAC but never
loaded.
8 instructions total
5 instructions in the Loop {LDD to LOOP1}
Smaller than any of the ISAs mentioned in the paper.
#0 in VEC is compiler telling HW that the loop can be executed as wide as HW has Lanes
A question. I certainly see the advantages of #0, and I think that #1 >>>>>> may be useful if you are doing a reduction and care about say overflow >>>>>> of intermediate values. Are there use use cases for values other than >>>>>> zero and one?
# not zero is how the compiler indicates a limit to the width of execution.
So, when a loop has a 3-iteration loop-carried dependency, compiler uses #3;
so, the recurrence is easily processed.
for( unsigned i=4; i<MAX; i++ )
a[i] = a[i] + b[i]*a[i-3];
This still vVM vectorizes but the width of iteration is limited to 3 so >>>>> that successive loops have ready data to use as operands.
{I am not saying this well (or in terminology of compiler writers)} >>>>
Well #0..#31 are available. -:)
AND we are not really thinking about #k >= 8 (power more than area)
So, first order::
a) Cray-vector machines will not vectorize.
b) SIMD-vector machines will not vectorize, either.
c) Both run out of vRF at r[-17].
d) so little expected damage to competitive perf.
What does the vector width value do to the uArch
Nothing, it has to do with loop-recurrences with is architectural.
I'm confused as you said earlier "#0 in VEC is compiler telling HW
that the loop can be executed as wide as HW has Lanes".
(eg does it control the scheduler or a uOp packer)?
Why would someone ever have a value other than #0 (autoconfig)?
for( int i = 3; i < MAX; i++ )
a[i] = x*b[i] + a[i-3];
a[3] is dependent on a[0]
a[4] is dependent of a[1]
...
Any vector element overlap is handled by store-to-load forwarding
(which is why VVM doesn't need the complex alias analysis
of SIMD or its permute/shuffle lane editing instructions).
The old paper https://www.sigarch.org/simd-instructions-considered-harmful/ compares SIMD
implementations with Vector implementations--as-if they were the only 2 games in town. The
paper contains the <toy> benchmark::
void daxpy(const size_t n, const double a, const double x[], double y[])
{
for (size_t i = 0; i < n; i++) {
y[i] = a*x[i] + y[i];
}
}
and goes on to extoll the beneficial properties of Vectors over SIMD using RISC-V as
exemplary. Let me pivot the conversation back to My 66000 virtual-Vector-Method vVM
(an alternative to both SIMD and Vectors). The above code compiles to::
daxpy:
MOV R5,#0
VEC #0,{}
LDD R6,[R3,R5] // LOOP transfers control back to here
LDD R7,[R4,R5]
FMAC R8,R2,R5,R6
ST R8,[R4,R5]
LOOP1 LT,R5,#1,R1
RET
On 9/26/2026 4:31 PM, MitchAlsup wrote:
The well published X-zillion 4KB random reads per second does not have
any of those 4KB pages touched by CPU instructions ?? So, it measures
only the PCIe perf and not memory perf.
Yeah. I can't really help but feel at least a little skeptical about
GB/s speeds being reported off of NVMe M.2 SSDs.
Maybe if there is some special mechanism to move blocks between the SSD
and RAM without needing to go through the CPU, or somehow multiple CPU
cores working in parallel for SSD IO.
BGB <cr88192@gmail.com> writes:
On 9/26/2026 4:31 PM, MitchAlsup wrote:
The well published X-zillion 4KB random reads per second does not have
any of those 4KB pages touched by CPU instructions ?? So, it measures
only the PCIe perf and not memory perf.
Yeah. I can't really help but feel at least a little skeptical about
GB/s speeds being reported off of NVMe M.2 SSDs.
Maybe if there is some special mechanism to move blocks between the SSD
and RAM without needing to go through the CPU, or somehow multiple CPU
cores working in parallel for SSD IO.
Perhaps you're thinking of the six-plus decade old technique of
Direct Memory Access (DMA) by peripherals?
Given modern interconnect (cpu and device) technologies (e.g. PCIe) pretty much _all_ peripheral data is moved without CPU intervention.
On 9/28/2026 10:29 AM, Scott Lurndal wrote:
BGB <cr88192@gmail.com> writes:
On 9/26/2026 4:31 PM, MitchAlsup wrote:
Either way, quite different from how HDDs were once done on x86:
Wait for IRQ;
Do a bunch of IN/OUT instructions to move the data via IO ports.
EricP <ThatWouldBeTelling@thevillage.com> posted:
(eg does it control the scheduler or a uOp packer)?
Why would someone ever have a value other than #0 (autoconfig)?
for( int i = 3; i < MAX; i++ )
a[i] = x*b[i] + a[i-3];
a[3] is dependent on a[0]
a[4] is dependent of a[1]
...
Any vector element overlap is handled by store-to-load forwarding
(which is why VVM doesn't need the complex alias analysis
of SIMD or its permute/shuffle lane editing instructions).
One could write:
for( int i = 3; i < MAX; i++ )
{
d3 = d2;
d2 = d1;
d1 = a[i] = x*b[i] + d3;
}
And change the ST->LD into R->R->R
Well, for me, it is like the weirdness of seeing videos where people
claim read/write speeds on NVMe SSD's that seem to encroach on the
territory I see when benchmarking "memcpy()" style operations.
Well, and much faster than the 300 MB/s I see with a SATA SSD (where
both the SSD and MOBO support SATA3, which theoretically has a higher
limit, but in testing I see 300MB/s, or the SATA2 limit).
BGB <cr88192@gmail.com> wrote:
Well, for me, it is like the weirdness of seeing videos where people
claim read/write speeds on NVMe SSD's that seem to encroach on the
territory I see when benchmarking "memcpy()" style operations.
As others noted the CPU generally don't touch the data for that kind
of test, it's basically a PCIe bandwidth & overhead test if the SSD is
fast enough.
One thing to remember is that for write the quoted speed won't be
sustained for very long unless you're buying extremely expensive
enterprise SSDs - most SSDs has a pseudo-SLC mode cache and then
"folds" the data to TLC or QLC.
Tom's Hardware always test SSDs "properly" (IMHO), so lets look at a
review of random semi-recent (Dec 2025) high-end consumer SSD, the
Corsair MP700 Pro XT [1].
For the 2TB model Corsair lists up to 14900MB/s read and 14500MB/s
write and someone mentioned it being capable of 3.3M IOPS but I can't
find a quote for that.
The ones Tom's lists is a bit lower at 12+GB/s and 2M IOPS but I have
no doubt Corsair really did manage to get the numbers they advertise.
But the section that I always also check is the "Sustained Write
Performance and Cache Recovery".
As usual for SSDs the sustained write performance is COMPLICATED - it
pretty much always is unless the interface is very limiting (IE SATA).
In this case the 2TB drive ran in pSLC mode at 13.7GB/s for 16 seconds (~220GB, using ~660GB of TLC flash cells), then switched to direct
write to the TLC flash at 4.2GB/s but there's intermittent short
period where the performance falls down to 1.96GB/s, showing the SSD
has to do "folding" (convert the pSLC data to native TLC) to free
flash storage.
The pSLC mode is there to make most normal operations lightning fast,
in most cases it'll have plenty of time to run the folding in the
background so it's not noticeable.
Most SSDs used dynamic pSLC cache using some (or all) of the free
flash, so the max cache size will shrink if the SSD is full.
I'm rather impressed by those numbers but then it's very much not a
cheap drive.
Once you get into cheaper drives the sustained write performance often
tanks a LOT, glares at an old Samsung 870 QVO (SATA QLC from 2020)
which when full has a sustained write speed of 35-40MB/s outside the
now minute (full, remember) pSLC cache. And that's far from the worst
QLC drive I've seen.
Flash has gotten faster since then, when I checked a recent budget
PCIe 5.0 x4 QLC drive the sustained write speed was shown as ~200MB/s
which was about in my expected range.
The manufacture advertises 5.0GB/s read and 4.2GB/s write speed for
this drive and I have no doubt they're not lying about that, I'm sure
it hits that advertised write speed until the pSLC cache runs full.
And to be honest, for a lot of people it's probably go work fairly
well at least if they don't fill it above say 80-90%. Yes, there will slowdowns when installing large programs and it's also a LOT cheaper
than the premium TLC drive above.
I do consider it misleading to not list the sustained write speed but
I'm not listing the vendor because NO ONE list it except for high-end Enterprise SSDs.
Well, and much faster than the 300 MB/s I see with a SATA SSD (where
both the SSD and MOBO support SATA3, which theoretically has a higher
limit, but in testing I see 300MB/s, or the SATA2 limit).
I have plenty of older SATA MLC or TLC SSDs which will fill the
interface in either read or write basically forever, because the
native write speed to the flash significantly exceed the interface
speed so there was no reason to implement a pSLC cache, it's all
native writes.
In practice this translates to 550-560 MB/s read or ~520 MB/write. If
you're getting 300MB/s you're actually exceeding the capabilities of a
SATA II link (even if only by a bit) :-)
Not sure if anyone still manufacturer/sells "good" SATA SSDs?, the
market has kind of moved on.
1. https://www.tomshardware.com/pc-components/ssds/corsair-mp700-pro-xt-2tb-ssd-review
| Sysop: | Amessyroom |
|---|---|
| Location: | Fayetteville, NC |
| Users: | 74 |
| Nodes: | 6 (0 / 6) |
| Uptime: | 02:20:35 |
| Calls: | 1,194 |
| Files: | 1,353 |
| D/L today: |
2 files (1,590K bytes) |
| Messages: | 291,157 |