From Newsgroup: comp.arch
MitchAlsup <
user5857@newsgrouper.org.invalid> writes:
anton@mips.complang.tuwien.ac.at (Anton Ertl) posted:
John Levine <johnl@taugh.com> writes:
-------------------
OTOH, the store-to-wide-load forwarding issue that I have discussed in
this thread is directly connected to the auto-vectorization of several
narrow loads into a wide load. So in this case the transformation has
direct connection to the microarchitectural pitfall.
I think we need to make a distinction between vectorization and
SIMDization. Vectors are like CRAY (and RISC-v) while SIMD is
most everyone else.
Architecturally, vectors are an instance of SIMD. Is there any other
instance?
In a compiler, the conversion of scalar code to code that uses SIMD instructions is called vectorization.
What you are thinking about is microarchitecture. Does the SIMD
instruction work on every scalar separately, or does it work on
several scalars in parallel.
One interesting aspect is that early SSE2 implementations worked 8
bytes at a time, so if the scalars were binary64 floats, you would
classify these implementations as "vectors", while if the scalars were binary32, you would classify them as "SIMD".
The bubble-sort benchmark sorts 4-byte integers. With
auto-vectorization, the program loads two adjacent integers with one
8-byte load; if they are reordered, it stores two integers with an
8-byte store. The microarchitectural pitfall occurs on a load after
such a store, and it slows the program down, but the program behaves
correctly, so the compiler works correctly and the architecture is
implemented correctly.
Maybe some future microarchitectures will include a way to avoid this
pitfall: E.g., it could notice that a given load sees relatively many
overlaps, and have a mechanism that splits it into sub-loads and merge
the results, basically undoing the vectorization. Would you classify
such a microarchitecture as vector or as SIMD?
One architectural difference between the Cray-1 and modern
architectures is that the Cray-1 only has 8-byte data, whether shorts,
ints, or floats (not sure about doubles); and I think that's even the
case for chars in C. So the issue of dealing with 4-byte or 1-byte
lanes does not pose itself in microarchitectures for this
architecture.
It does pose itself for RISC-V. I have not looked at RISC-V vectors
and their implementations, but somehow I doubt that all of them will
perform string loads one byte at a time.
Reading up about the NEC SX-Aurora <
https://en.wikipedia.org/wiki/SX-Aurora_TSUBASA>, the spiritual
successor of the Cray-1, it has 64 logical vector registers with
256x64 bits length. It's implementation uses "32-fold parallel SIMD
units" (obviously using SIMD in your sense, not the architectural
sense; I would rather call it "32x64-bit-wide functional units"), and
with three FMA units this results in up to 192 DP FP ops per cycle. I
expect that it has a similar performance problem if a load partially
overlaps a store, especially if the distance is not a multiple of the
FU width.
What would a 2018-vintag Cray-1 successor look like. Probably much
like the NEC SX-Aurora, i.e., it would also have this problem.
Given that we have not seen an SX-Aurora successor in 8 years, despite
new types of HBM appearing and the demand for SIMD (with small to very
small components) being huge, it seems like NEC has finished their SX
line with the Aurora.
In modern machines LDs are well pipelined, often multi-lane; so,
when a compiler auto-SIMDs a gaggle of LDs into a LDSIMD, one
then needs a similar means to extract the gaggle from the SIMD.
This extract is less likely to have the same properties of the
LD pipeline.
Cartainly even the load-only version of bubble-sort performs slower on
Rocket Lake <
2026Sep21.171615@mips.complang.tuwien.ac.at>. OTOH, microarchitectures tend to have fewer load units than ALU units, so
even if the scalar values are needed, one probably can construct a
cases where a wide load followed by extraction and processing is
faster than two narrow loads followed by processing. In general,
though, it probably is only beneficial if the data then does not need
to be extracted, but the further processing can also happen in a SIMD
way.
E.g., I discuss [ertl24] a number of microarchitectural
pitfalls that many Forth systems fall into, not because of
sophistication, but because the Forth systems use techniques that
worked fine in the microprocessors of the 1970s and 1980s, and
actually still work fine on threaded-code systems on modern CPUs, but
fall into these pitfalls in native-code systems.
Is this related to the over aggressive branch predictors of modern ?
Either branches themselves or indirect calls.
One of the pitfalls is related to branch prediction: Using a call
instruction and then not returning, but popping the return address off
the stack to get the position of the call instruction as data; this
results in the hardware return stack getting out-of-sync and results
in all returns at the current level and further out to be
mispredicted.
What do you mean by "over aggressive branch prediction"? Either a
branch predictor predicts correctly or it mispredicts. In what case
would it be "over aggressive"?
- anton
--
'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
Mitch Alsup, <
c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
--- Synchronet 3.22a-Linux NewsLink 1.2