From Newsgroup: comp.arch
On 9/14/2026 12:49 PM, MitchAlsup wrote:
BGB <cr88192@gmail.com> posted:
On 9/12/2026 1:31 PM, John Dallman wrote:
In article <6aa588d4.1140437@news.eternal-september.org>,
quadibloc@invalid.com (John Savard) wrote:
--------------
Meanwhile... I am still having pretty good results with 64 unified GPRs.
The move from XG2 to XG3 had cost a few usable registers:
R0..R4 have effectively left the GPR space...
Zero, LR/RA, SP, and GP, TP.
In My 66000::
R0 receives the return address on a call, but otherwise is just
another GPR. In safe-stack mode, R0 is just another GPR.
ENTER and EXIT use SP (R31) implicitly, otherwise SP would be
Just another GPR.
FP is Just another GPR
Do to 64-bit Displacements, there is no need for Global pointer.
When needed Thread Local Pointer becomes R16.
Part of the register change-over was due mostly to the RV64G merger...
XG3, in merging with RV64G, adopted the same register space as RV64G.
It also mostly adopted a variant of RV64G's C ABI.
Though not quite the same as either the LP64 or LP64D ABI, but sort of a hybrid;
And, I ended up tweaking the rules further.
There is the XG3N ABI, which has 16 arguments and a spill space, but
this is not currently the default ABI.
Though, ATM, this was more a thing of "reducing friction".
Can note:
My runtime:
TP is used for the TaskInfo structure
Similar to the TEB on Windows;
On X64, would have been encoded with an SEG_FS prefix.
GP is used for the base of the ".data" section.
Neither PC-rel or Abs64 would be valid in the PBO rules.
RV+GCC uses TP and GP differently.
--------
Likewise...
Not too long ago, saw something where someone was (once again) pushing
for bank-switched registers, and hardware context switching, in the RV
interrupt handling mechanism.
One of the benefits of My 66000 having all thread data "effectively"
reside in memory, is that a simulator can simply read/write simulator
memory, a small implementation can use a single set of control registers
and a register file loading and storing like a write back cache, while
a GBOoO machine can support the model with bank switching. The running software cannot tell the difference.
OK.
In the TaskInfo structure, there is a space for saved registers.
Generally, ISR's like SYSCALL and TLBMISS save the registers there,
rather than on the ISR stack, mostly because this avoids needing to
handle them all twice during a context switch.
Does mean that the TaskInfo needs to be set up before these ISRs can be
used. Generally, the kernel needs to set up the TaskInfo structure for
the kernel, then spawn the SYSCALL task, before it can bring up virtual memory.
This is partly backwards from x86, where one would bring up the MMU first.
Meanwhile, naive and simple interrupt handling tends to work out
cheapest and fastest in practice...
The most important thing to remember here, is to stay as far away
from x86 as possible. ...
Errm, I don't think anyone here is proposing bringing back the IDT, GDT,
and TSS.
But, yeah, the interrupt mechanism I have is pretty close to the minimum:
Save off the needed state into a few CRs;
Set up interrupt mode;
Which swaps the SP and SSP registers in decoding;
Generate input entry point address from a base-vector via bit-slicing.
You don't really want the effective register file to be bigger than it
needs to be. Likewise, it is not like there is any good way to make a
register/memory side-channel that is faster than the main register <->
memory path (via load/store instructions).
My 66000, due to the way Thread.State is managed, can start loading
in new state before it stores back current state. Software has to
store a register before it can load a register. Hardware can load
a buffer, and then read the old out to another buffer, while writing
the new in, and then finally migrate the old buffer to memory. Saving
a round trip latency to wherever the new data is coming from {L2, L3,
DRAM}.
OK.
Probably requires something a little more advanced than the
direct-mapped cache I am currently using to be effective.
Though, to make use of loading and saving a context at the same time,
would mean needing to transfer control without the control first being intercepted by an interrupt handler (meaning the decision to perform a
context switch would need to itself be driven by hardware).
Well, unless the ISR is itself another context switch, meaning two such context switches per interrupt...
At present, the interrupts' context effectively disintegrates/disappears
when it returns from an interrupt.
And, if you had an interrupt mechanism that basically just invokes blobs
of hidden firmware for the interrupt dispatch (and return), this could
be done, but doesn't gain anything.
Even I discarded having the control-transfer-switch logic do anything
other than write-back-cache above. This also means SW can place what
ever kinds of tables it wants, wherever it wants, of whatever size it
wants without any page boundaries or alignments getting in the way.
OK.
------------
So an instruction word starting with 11 but not 1111 is composed of
11 followed by five six-bit prefix fields.
Will you object if the compiler writer decides not to use this part of
the ISA?
This is where I draw the line, if the compiler can't use it, it does not
get in--except for reading and writing of control registers.
Mostly similar.
I mostly avoid features that would require hand-written ASM or similar
to use.
Except for some very niche instructions, like color-cell encoding
helpers and similar. This being mostly because this is something that
can happen a lot in some use cases, and needs to be fast.
In the TestKern GUI, I ended up with a strategy though, of first using a
quick and dirty ISA driven color-cell encoder, and then if the blocks
had previously used this one, and haven't changed for a while, it goes
back over them and re-encodes them with a higher quality but slower SW
based color-cell encoder.
Fast strategy:
Select the min and max colors:
Internally shuffles the RGB555 bits around and does compares.
Then, for each block of 4 pixels,
map its indices between the min and max.
So, for the RGB555 -> Luma:
0rrrrrgggggbbbbb => ggrbgrbg (starting at HOBs of each component)
Generate values, and a comparison matrix, and use this to select.
To map indices:
Ymin = rgbtoluma(Cmin);
Ymax = rgbtoluma(Cmax);
Ymid = (Ymin + Ymax)>>1;
Ymlo = (Ymin + Ymid)>>1;
Ymhi = (Ymid + Ymax)>>1;
Yc0 = rgbtoluma(Clr0);
Yc1 = ...
Ix0 = (Yc0 < Ymid) ?
((Yc0 < Ymlo) ? 00 : 01) :
((Yc0 < Ymhi) ? 10 : 11) ;
Ix1 = ...
Ix2 = ...
Ix3 = ...
Then run this logic for each row of pixels.
In the software encoder, one way to boost quality is to to proper luma
math, and to select among several possible color axes (typically Cyan/Magenta/Yellow/White), choosing whichever axis has the highest
contrast.
Say:
Ycy = (4*Cb+3*Cg+1*Cr)/8;
Ymg = (4*Cr+3*Cb+1*Cg)/8;
Yye = (4*Cg+3*Cr+1*Cb)/8;
Ywh = (2*cg+1*cr+1*cb)/4;
Arguably, a general purpose CPU doesn't need this stuff in hardware though.
So:
LDTEX
CCENC
...
Remain as mostly super-niche stuff that only really remains because
certain use-cases need this stuff to be fast.
Nevermind whether or not the GUI mode is highly used.
Decided to leave out a detour about color-cell based video codecs (also
a bit niche, but for my own uses I was primarily using color-cell based designs rather than MPEG derived designs).
Well, ironically, apart from my UPIC image format, which ironically is
an offshoot of my experiments with trying to make MPEG like codecs fast...
And, UPIC then is basically "What if we took T.81 JPEG and used Rice
Coding and Block-Haar and the RCT transform...". Ironically, compression
is still pretty competitive with T.81 JPEG, but is a little faster. It
also supports a PNG like lossless mode, which ironically for many images
both beats PNG both on compression and decode speeds.
The long-standing issue of color-cell designs being basically, how to
most efficiently encode the color endpoints and block patterns.
Well, more so in the absence of entropy encoding, because the magic of
entropy coding coming with the penalty of making everything slow.
He keeps going at it...
And we keep humoring him--sad to see is rate of forward progress has
not changed in 5 years...
Not really sure the merit of endlessly bashing at non-sane encoding schemes.
He fails to see he is barking up a large bush instead of a tree.
Yeah.
There is a difference between picking a design that basically works.
And thrashing around with stuff that doesn't.
In my case, I mostly just seem to be running low on stuff that is
"actually interesting".
In my case, I eventually realized I couldn't have *everything* I wanted
all at the same time, so had dropped off the lower priority items.
Architecture is as much about what you leave out as what you let in.
I eventually ended up dropping 16-bit ops, but kept 64 GPRs and predication.
This was the direction that maximized speed.
OTOH, 16-bit ops can make sense if the goal is to instead maximizing
code density.
But, they are not entirely exclusive:
XG3 also has OK code density because keeping instruction counts small
also happens to reduce binary size (even if not the primary goal).
RV64GC isn't horribly slow, mostly for sake of the 32-bit instructions.
At one point, XG3 moved into the lead on code density, but RV64GC moved
back into the lead when I added some experimental 32-bit encodings for:
LW Xd, Disp16u*4(GP) //256K
LD Xd, Disp16u*8(GP) //512K
Both still beating the original form of RV64GC though on code density.
Though, as some ELF PIE binaries show, there could be a size advantage
to using a dynamically linked C library, whereas at present
BGBCC/TestKern is using a static linked C library.
Otherwise, at least getting closer on trying to get "ld.so" working:
Sorted out some issues that were breaking the VFS syscalls;
Also now have mmap seemingly working (*1).
At present, "ld.so" fails with an assert in its "init_tls" stage.
*1: It is a little funky ATM as it is needing to pretend to implement a
4K page size when the underlying CPU isn't actually using 4K pages.
Seemingly the ELF loader then tries to create multiple overlapping
mmaps, but the current implementation effectively merges the mmaps
together in this case.
Well, because apparently this stuff is partly hard-coded to assume a 4K
page size and doesn't really work if the page size is different. But,
does sort of work if the implementation pretends to support 4K pages.
...
--- Synchronet 3.22a-Linux NewsLink 1.2