• Achieving the Impossible with 31 Registers

    From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sat Sep 12 18:11:05 2026
    From Newsgroup: comp.arch

    At present, I have defined on my site Concertina II with block
    structure, and Concertina III without. Recently, I had come up with a
    scheme for Concertina III that made it look less messy.

    Concertina III thus has variable-length instructions, forcing
    instruction decode to be somewhat sequential.

    I had let myself be persuaded that block structure is bad, since
    agreement on that was nearly unanimous on this board. Yet I really
    wanted to have parallel decoding - as if every instruction was only 32
    bits long - and yet have an instruction set with instructions of
    different lengths.

    The one who wastes time trying to achieve the impossible is a fool.
    The one who can achieve what only seemed impossible has contributed to
    human progress.

    Well, there certainly _is_ a way to do this.

    At present, due to my vain efforts at developing a Concertina II or
    higher ISA, I am at the point where I can fit a complete RISC+ (RISC
    with base-index addressing added) instruction set into 3/4 of the
    opcode space for 32-bit instructions.

    So I have all the 32-bit values starting with the two bits 11 to play
    with.

    I could divide the possibilities into two groups; one contains several prefixes, and the other contains instructions which take the prefixes.

    So now I can lengthen a 32-bit instruction anywhere without a block
    structure. That's all well and good, but then to avoid a mess, I
    probably have to stipulate that the prefixes and the things prefixed
    must be in the same basic block. And I've got some exotic internal
    register that's added to what has to be saved in an interrupt.

    This is where I had been for a while until recently.

    In my ISA, I have banks of 32 registers, one for integers, one for floating-point numbers. The integer registers are divided up into
    groups with special purposes:

    0: by convention, used for the return value of an integer-valued
    function
    1: by convention, used for pointer values returned by functions with
    values that can't fit in a simple register
    1-7: may be used as index registers
    9-15: may be used as base registers for 20-bit displacements
    16: may be used as the implicit base register for 15-bit displacements
    17-23: may be used as base registers for 12-bit displacements
    24: points to an array of pointers for Array Mode, which allows a
    program to have multiple large arrays that are simply addressed but
    don't fit into 64K without having to assign a scarce base register to
    each one
    25-31: may be used as base registers for 16-bit displacements

    I think I may have had the roles of registers 16 and 24 switched
    around, but since 15-bit and 12-bit addressing are closely associated,
    this seems a better assignment.

    Well, now I've found a use for integer register 8.

    When a 32-bit word, the first two bits are 11, and the next two bits
    of which are not both 1, is encountered during execution, it is
    executed as an instruction which does the following:

    Integer register 8 is shifted right by 30 bits.
    The least significant 30 (not 28) bits of the instruction word are
    loaded into the first 30 bits of integer register 8.

    When a 32-bit word starting with 1111 is encountered, the most
    significant six bits in integer register 8 are used to determine what
    to do with that word... and integer register 8 is shifted left 6 bits.

    Under some circumstances, taking the most significant 6 bits of
    integer register 8 and shifting left 6 bits may be repeated once
    before the instruction including the remaining 28 bits of the word
    starting with 1111 is executed.

    This way, the instruction prefixes are stored in a visible register,
    that is saved and restored normally during interrupts. And if you
    don't use the fancy instructions that are added to the ISA by this
    mechanism, then you have all 32 integer registers available to your
    program.

    So an instruction word starting with 11 but not 1111 is composed of 11
    followed by five six-bit prefix fields. What is their form?

    In each six bit prefix field, the first two bits (which, at least in
    the first position, can't be 11) are special.

    They can be:

    00 - indicates a four-bit prefix which turns 28 bits into a 32-bit
    instruction from a secondary auxilliary instruction set
    01 - indicates a four-bit prefix which turns 28 bits into 32 bits that
    are, or form part of, a non-executable immediate value used by some
    other instruction - or something else that isn't the start of an
    executable instruction, such as the second half of a 64-bit
    instruction
    10 - indicates a pair of two-bit prefixes, which are prefixed to the
    two 16-bit halves of a 32-bit word

    10 is used to create code where the instructions can vary in length,
    being nominally (before you count the prefixes and overhead) 16 bits,
    32 bits, 48 bits, and so on.
    But since I insist on parallel decoding, those paired prefixes can't
    just apply to a non-prefixed 32-bit word. Instead, they will have to
    apply to a 1111-word prefixed by an 01-prefix (or at least something
    that acts like one).
    Plus, although the four-bit prefix has to be added to the 28 remaining
    bits first before you have 32 bits to split into two halves and add
    two bits of prefix to each half... the 01-prefix has to follow the
    10-prefix, since it's the 10-prefix that indicates that the 1111-word
    in question has to have *two* prefixes applied to it.

    As a bonus from doing it that way (necessary so that the prefixes will
    have the "prefix property"), the second prefix doesn't have to start
    with 01. Instead, those two bits can be an additional one-bit
    post-prefix.

    So a prefix of the form

    10abgh uvpqrs

    when applied to an instruction word of the form

    11xxxxxxxxxxxxxxyyyyyyyyyyyyyyyy

    produces two 19-bit elements in the instruction stream of the form

    ab upqrsxxxxxxxxxxxxxx gh vyyyyyyyyyyyyyyyy

    if this is less confusing than my description above.

    Now, at this point, we can see why the prefix-containing 11-word
    begins by shifting the existing contents of register 8 thirty bits to
    the right.
    We might want to compose a stretch of code consisting of
    variable-length instructions, which means a bunch of 32-bit 1111-words
    which each require a _pair_ of prefixes.
    But a 11-word with prefixes contains *five* prefixes, which is an odd
    number.
    So now we don't have to always waste a prefix in this case.
    We start with a 11-word, put five prefixes in it, like this: 11(1a)(1b)(2a)(2b)(5b)
    Follow it with two 1111-words; the first uses prefixes (1a) and (1b),
    and the second of which uses prefixes (2a) and (2b).
    Then we insert another 11-word with these five prefixes:
    11(3a)(3b)(4a)(4b)(5a)
    And now we can proceed with three 1111-words, which use...
    prefixes (3a) and (3b), then
    prefixes (4a) and (4b), and finally
    prefixes (5a) and (5b).

    The first two bits of a prefixed 19-bit element are interpreted as
    follows:

    00 - 18-bit short instruction starting with 0
    01 - 18-bit short instruction starting with 1
    10 - The start of a 33-bit or longer instruction
    11 - Not the start of an instruction

    In the event that 11 begins an immediate value, the post-prefix bit
    will go to waste.

    So with this perhaps insanely complex scheme - although I don't think
    of it as being as insane as my *last* notion - I feel I have achieved
    what, previously, I had only been able to achieve with block
    structure, but without having to impose a block structure.

    Note also that an instruction, including any immediates used by the instruction, cannot straddle instruction types. Instructions within
    the stream of 19-bit elements, and instructions within the stream of
    prefixed 32-bit words, are separate from each other. Each instruction
    executes when all parts of the instruction have been read from the
    instruction stream, so prefix types can be mixed arbitrarily; also, if
    a non-prefixed instruction takes an immediate value, that can be
    contained in a following 1111-word; thus, there are two disjoint
    instruction streams, not three; the one of 19-bit elements, and the
    one containing both non-prefixed instructions and prefixed 32-bit
    words.

    When I replace my existing Concertina II and Concertina III with a
    Concertina II based on this, rest assured there will be explanatory
    diagrams.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From jgd@jgd@cix.co.uk (John Dallman) to comp.arch on Sat Sep 12 19:32:40 2026
    From Newsgroup: comp.arch

    In article <6aa588d4.1140437@news.eternal-september.org>,
    quadibloc@invalid.com (John Savard) wrote:

    In my ISA, I have banks of 32 registers, one for integers, one for floating-point numbers.

    Are there instructions for converting integers to floating point and
    vice-versa without going via memory?

    24: points to an array of pointers for Array Mode, which allows a
    program to have multiple large arrays that are simply addressed but
    don't fit into 64K without having to assign a scarce base register
    to each one.

    Descriptors in memory are slow.

    So an instruction word starting with 11 but not 1111 is composed of
    11 followed by five six-bit prefix fields.

    Will you object if the compiler writer decides not to use this part of
    the ISA?

    John
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sat Sep 12 18:53:51 2026
    From Newsgroup: comp.arch

    On Sat, 12 Sep 2026 19:31 +0100 (BST), jgd@cix.co.uk (John Dallman)
    wrote:
    In article <6aa588d4.1140437@news.eternal-september.org>, >quadibloc@invalid.com (John Savard) wrote:

    In my ISA, I have banks of 32 registers, one for integers, one for
    floating-point numbers.

    Are there instructions for converting integers to floating point and >vice-versa without going via memory?

    It is definitely my intention to provide such instructions, yes.

    24: points to an array of pointers for Array Mode, which allows a
    program to have multiple large arrays that are simply addressed but
    don't fit into 64K without having to assign a scarce base register
    to each one.

    Descriptors in memory are slow.

    Just pointers.

    So an instruction word starting with 11 but not 1111 is composed of
    11 followed by five six-bit prefix fields.

    Will you object if the compiler writer decides not to use this part of
    the ISA?

    I don't feel it's my place to object. If the extra instructions in
    that part of the ISA contribute something to performance in a certain
    class of applications, then if that compiler serves those
    applications, it will be up to the compiler writer's customers to
    object.

    The basic RISC part of the ISA is complete in itself. One significant
    loss, though, is that without this fancy stuff, immediate values
    aren't available. Otherwise, one can do without it fairly well for
    many types of application - but the specialized instructions made
    available may be valuable in some cases.

    On second thought, though, I think that I will have to give up the
    extra post-prefix bit in order to let decoding be more parallel.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sat Sep 12 23:20:10 2026
    From Newsgroup: comp.arch

    I see I had made one fundamental error in the design.
    I should have tilted towards adding an extra bit to the words with
    most of the 32-bit words involved, because then, at the cost of one
    more bit there, I can shorten each header by one bit in the words with
    headers - and, as they contain multiple headers, multiple bits are
    gained.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Sun Sep 13 18:23:41 2026
    From Newsgroup: comp.arch

    On 9/12/2026 1:31 PM, John Dallman wrote:
    In article <6aa588d4.1140437@news.eternal-september.org>, quadibloc@invalid.com (John Savard) wrote:

    In my ISA, I have banks of 32 registers, one for integers, one for
    floating-point numbers.

    Are there instructions for converting integers to floating point and vice-versa without going via memory?


    Meanwhile... I am still having pretty good results with 64 unified GPRs.


    The move from XG2 to XG3 had cost a few usable registers:
    R0..R4 have effectively left the GPR space...
    Zero, LR/RA, SP, and GP, TP.

    And, the register balance is currently:
    28 callee save; 31 scratch; 5 special.
    Vs in XG2:
    31 callee save, 30 scratch; 3 special.

    Well, and the RV64 ABI's register assignments are more ad-hoc.

    I ended up deviating from the original ABIs to get 28, but 28 put it
    closer to balance, and BGBCC seems to work best when the number of
    callee save and scratch registers is roughly in balance.

    Must resist temptation to go into big detours about ABI design...



    24: points to an array of pointers for Array Mode, which allows a
    program to have multiple large arrays that are simply addressed but
    don't fit into 64K without having to assign a scarce base register
    to each one.

    Descriptors in memory are slow.


    Yes.


    Likewise...

    Not too long ago, saw something where someone was (once again) pushing
    for bank-switched registers, and hardware context switching, in the RV interrupt handling mechanism.

    Meanwhile, naive and simple interrupt handling tends to work out
    cheapest and fastest in practice...

    You don't really want the effective register file to be bigger than it
    needs to be. Likewise, it is not like there is any good way to make a register/memory side-channel that is faster than the main register <->
    memory path (via load/store instructions).

    And, if you had an interrupt mechanism that basically just invokes blobs
    of hidden firmware for the interrupt dispatch (and return), this could
    be done, but doesn't gain anything.


    When I did look into/experiment with bank switched registers, it quickly became obvious that while the mechanism could work, it would have ended
    up around 5x slower than the existing mechanism. Replacing a pipelined load/store sequence, with stalling the pipeline for up to 15 clock
    cycles for every instruction that accesses registers that need to be
    swapped out (as going entirely over to BRAM for a 6R3W regfile is not
    viable).


    Ironically, the mechanism could exist as a possible way to support
    something like RV-V, but it is losing out to to "don't bother, just fake
    it in software if it comes up" (where I can continue using a SIMD design
    that isn't quite so horribly expensive).

    Granted, this stuff could maybe make sense on hardware where one can
    afford to have like 32 or 64 kilobits in the register file...



    So an instruction word starting with 11 but not 1111 is composed of
    11 followed by five six-bit prefix fields.

    Will you object if the compiler writer decides not to use this part of
    the ISA?


    He keeps going at it...

    Not really sure the merit of endlessly bashing at non-sane encoding schemes.

    In my case, I eventually realized I couldn't have *everything* I wanted
    all at the same time, so had dropped off the lower priority items.



    In my case, I had recently made some progress towards getting RV64GC ELF
    PIE + glibc binaries working on TestKern:
    Well, turns out the hidden LR tagging issue was a major issue blocking
    my past attempts (as is, I have moved away from using LR tagging in
    RV64GC mode).

    Next issue it seems, is getting the Linux style system calls working
    well enough that ld.so and glibc are happy with them.


    Like, the slightly confusing mess that is the "newfstatat" system call,
    which is seemingly necessary to get working however ld.so wants, or it
    refuses to work. Well, and it is going to need to be necessary to
    support file mappings and similar (glibc seemingly wants to use mmap to
    pull in SO contents).

    Current poor man's solution is to just sort of have the mmap call
    allocate the buffer and read in the contents though (rather than create
    a CoW mapping).

    The kernel level interfacing for some of this stuff is annoyingly poorly documented (and, the closest thing I have to "debugability" is that at
    least now "ld.so" prints out 'helpful' error messages before it crashes
    out).

    The glibc library code is not exactly easy to decipher.

    ...


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Mon Sep 14 04:36:52 2026
    From Newsgroup: comp.arch

    On Sun, 13 Sep 2026 18:23:41 -0500, BGB <cr88192@gmail.com> wrote:

    Not really sure the merit of endlessly bashing at non-sane encoding schemes.

    In my case, I eventually realized I couldn't have *everything* I wanted
    all at the same time, so had dropped off the lower priority items.

    I have to admit that this is a valid way of characterizing what I've
    been up to.
    Managing to squeeze base-index addressing with 16-bit displacements
    with RISC-size 32-register banks ought to be enough of an achievement.
    But the wild and wacky encoding schemes I've been coming up with had
    one simple goal:
    - gain the efficiency of potentially parallel instruction decoding, as
    is possible if all instructions are the same length;
    - but at the same time, have instructions that differ in length;
    - and which include in-line immediate values, which militates against
    schemes that put special bits in front of each word to flag the words
    that aren't to be decoded as the start of an instruction.

    And now I have the new constraint of trying to do that without block
    structure.

    At first, I thought that without block structure, it just couldn't be
    done.
    Of course, it could be done this way: flag 32-bit words with a first
    few bits that indicate they're special - and precede those words with
    another 32-bit word that contains several small headers, that both
    rebuild those words to a full 32 bits, and indicate if they're the
    beginning of an instruction or not.
    But this had serious issues. Between the header word, and the
    instruction words it affected, there could be no transfers of control.
    And the header word becomes an extra part of the machine state.
    This is what makes it impractical.
    So I've found a way around that - by dedicating one register to
    processing the specialized instructions, if they are used. This tames
    the aligned header scheme.
    It's still not a total solution. Because the headers are going through
    a register, now decoding is tangled up with execution, and thus
    rendered more serial. As long as something weird like directly
    changing the register isn't being done, this may just mean a later
    parallel decoding layer rather than forcing serial decoding.
    But at least what the headers do is clearly definable and unambiguous,
    and extending the machine state is avoided.

    Why do I waste time on such schemes? Well, the world already has
    plenty of ordinary computer ISAs. I want to achieve the extraordinary,
    even though it appears impossible, and is therefore difficult and
    complicated.
    Can I improve it? Can I get rid of the rough edges so as to come up
    with something that doesn't seem so wild as to be insane? That may
    take iterations.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 14 17:51:09 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:
    -------------
    Can I improve it?

    Yes, get a compiler up and running so you can read the resulting assembly
    code. Then you are in a position to measure how well things are going (or not).

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 14 17:49:28 2026
    From Newsgroup: comp.arch


    BGB <cr88192@gmail.com> posted:

    On 9/12/2026 1:31 PM, John Dallman wrote:
    In article <6aa588d4.1140437@news.eternal-september.org>, quadibloc@invalid.com (John Savard) wrote:
    --------------

    Meanwhile... I am still having pretty good results with 64 unified GPRs.


    The move from XG2 to XG3 had cost a few usable registers:
    R0..R4 have effectively left the GPR space...
    Zero, LR/RA, SP, and GP, TP.

    In My 66000::

    R0 receives the return address on a call, but otherwise is just
    another GPR. In safe-stack mode, R0 is just another GPR.

    ENTER and EXIT use SP (R31) implicitly, otherwise SP would be
    Just another GPR.

    FP is Just another GPR

    Do to 64-bit Displacements, there is no need for Global pointer.

    When needed Thread Local Pointer becomes R16.
    --------
    Likewise...

    Not too long ago, saw something where someone was (once again) pushing
    for bank-switched registers, and hardware context switching, in the RV interrupt handling mechanism.

    One of the benefits of My 66000 having all thread data "effectively"
    reside in memory, is that a simulator can simply read/write simulator
    memory, a small implementation can use a single set of control registers
    and a register file loading and storing like a write back cache, while
    a GBOoO machine can support the model with bank switching. The running
    software cannot tell the difference.

    Meanwhile, naive and simple interrupt handling tends to work out
    cheapest and fastest in practice...

    The most important thing to remember here, is to stay as far away
    from x86 as possible. ...

    You don't really want the effective register file to be bigger than it
    needs to be. Likewise, it is not like there is any good way to make a register/memory side-channel that is faster than the main register <-> memory path (via load/store instructions).

    My 66000, due to the way Thread.State is managed, can start loading
    in new state before it stores back current state. Software has to
    store a register before it can load a register. Hardware can load
    a buffer, and then read the old out to another buffer, while writing
    the new in, and then finally migrate the old buffer to memory. Saving
    a round trip latency to wherever the new data is coming from {L2, L3,
    DRAM}.

    And, if you had an interrupt mechanism that basically just invokes blobs
    of hidden firmware for the interrupt dispatch (and return), this could
    be done, but doesn't gain anything.

    Even I discarded having the control-transfer-switch logic do anything
    other than write-back-cache above. This also means SW can place what
    ever kinds of tables it wants, wherever it wants, of whatever size it
    wants without any page boundaries or alignments getting in the way. ------------
    So an instruction word starting with 11 but not 1111 is composed of
    11 followed by five six-bit prefix fields.

    Will you object if the compiler writer decides not to use this part of
    the ISA?

    This is where I draw the line, if the compiler can't use it, it does not
    get in--except for reading and writing of control registers.

    He keeps going at it...

    And we keep humoring him--sad to see is rate of forward progress has
    not changed in 5 years...

    Not really sure the merit of endlessly bashing at non-sane encoding schemes.

    He fails to see he is barking up a large bush instead of a tree.

    In my case, I eventually realized I couldn't have *everything* I wanted
    all at the same time, so had dropped off the lower priority items.

    Architecture is as much about what you leave out as what you let in.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 14 19:16:12 2026
    From Newsgroup: comp.arch

    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    BGB <cr88192@gmail.com> posted:

    On 9/12/2026 1:31 PM, John Dallman wrote:
    In article <6aa588d4.1140437@news.eternal-september.org>,
    quadibloc@invalid.com (John Savard) wrote:
    --------------

    Meanwhile... I am still having pretty good results with 64 unified GPRs.


    The move from XG2 to XG3 had cost a few usable registers:
    R0..R4 have effectively left the GPR space...
    Zero, LR/RA, SP, and GP, TP.

    In My 66000::

    R0 receives the return address on a call, but otherwise is just
    another GPR. In safe-stack mode, R0 is just another GPR.

    It deserves its own register class in a compiler because it acts
    as a substitute for IP as a base register, as "no register" as
    index register, as thread-defined rounding mode in CVTxx, as the
    address to jump back to with "ret" in the absence of a safe stack.

    So it is a bit special.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Brian G. Lucas@bagel99@gmail.com to comp.arch on Tue Sep 15 13:05:50 2026
    From Newsgroup: comp.arch

    On 9/14/26 2:16 PM, Thomas Koenig wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    BGB <cr88192@gmail.com> posted:

    On 9/12/2026 1:31 PM, John Dallman wrote:
    In article <6aa588d4.1140437@news.eternal-september.org>,
    quadibloc@invalid.com (John Savard) wrote:
    --------------

    Meanwhile... I am still having pretty good results with 64 unified GPRs. >>>

    The move from XG2 to XG3 had cost a few usable registers:
    R0..R4 have effectively left the GPR space...
    Zero, LR/RA, SP, and GP, TP.

    In My 66000::

    R0 receives the return address on a call, but otherwise is just
    another GPR. In safe-stack mode, R0 is just another GPR.

    It deserves its own register class in a compiler because it acts
    as a substitute for IP as a base register, as "no register" as
    index register, as thread-defined rounding mode in CVTxx, as the
    address to jump back to with "ret" in the absence of a safe stack.

    So it is a bit special.

    The current compiler implementer was lazy and just never allocated R0.
    This could be revisited if register pressure was really an issue.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Tue Sep 15 19:54:57 2026
    From Newsgroup: comp.arch


    "Brian G. Lucas" <bagel99@gmail.com> posted:

    On 9/14/26 2:16 PM, Thomas Koenig wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    BGB <cr88192@gmail.com> posted:

    On 9/12/2026 1:31 PM, John Dallman wrote:
    In article <6aa588d4.1140437@news.eternal-september.org>,
    quadibloc@invalid.com (John Savard) wrote:
    --------------

    Meanwhile... I am still having pretty good results with 64 unified GPRs. >>>

    The move from XG2 to XG3 had cost a few usable registers:
    R0..R4 have effectively left the GPR space...
    Zero, LR/RA, SP, and GP, TP.

    In My 66000::

    R0 receives the return address on a call, but otherwise is just
    another GPR. In safe-stack mode, R0 is just another GPR.

    It deserves its own register class in a compiler because it acts
    as a substitute for IP as a base register, as "no register" as
    index register, as thread-defined rounding mode in CVTxx, as the
    address to jump back to with "ret" in the absence of a safe stack.

    So it is a bit special.

    The current compiler implementer was lazy and just never allocated R0.
    This could be revisited if register pressure was really an issue.

    But, since in general, all 32 GPRs can hold any value desired
    interchangeably; it maters less than when some registers have
    defined values (R0) or defined support functions (ld.so) and
    the 32 register name space only contains effective 27-28 registers.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BGB@cr88192@gmail.com to comp.arch on Tue Sep 15 17:46:44 2026
    From Newsgroup: comp.arch

    On 9/14/2026 12:49 PM, MitchAlsup wrote:

    BGB <cr88192@gmail.com> posted:

    On 9/12/2026 1:31 PM, John Dallman wrote:
    In article <6aa588d4.1140437@news.eternal-september.org>,
    quadibloc@invalid.com (John Savard) wrote:
    --------------

    Meanwhile... I am still having pretty good results with 64 unified GPRs.


    The move from XG2 to XG3 had cost a few usable registers:
    R0..R4 have effectively left the GPR space...
    Zero, LR/RA, SP, and GP, TP.

    In My 66000::

    R0 receives the return address on a call, but otherwise is just
    another GPR. In safe-stack mode, R0 is just another GPR.

    ENTER and EXIT use SP (R31) implicitly, otherwise SP would be
    Just another GPR.

    FP is Just another GPR

    Do to 64-bit Displacements, there is no need for Global pointer.

    When needed Thread Local Pointer becomes R16.


    Part of the register change-over was due mostly to the RV64G merger...
    XG3, in merging with RV64G, adopted the same register space as RV64G.


    It also mostly adopted a variant of RV64G's C ABI.
    Though not quite the same as either the LP64 or LP64D ABI, but sort of a hybrid;
    And, I ended up tweaking the rules further.


    There is the XG3N ABI, which has 16 arguments and a spill space, but
    this is not currently the default ABI.

    Though, ATM, this was more a thing of "reducing friction".


    Can note:
    My runtime:
    TP is used for the TaskInfo structure
    Similar to the TEB on Windows;
    On X64, would have been encoded with an SEG_FS prefix.
    GP is used for the base of the ".data" section.
    Neither PC-rel or Abs64 would be valid in the PBO rules.
    RV+GCC uses TP and GP differently.


    --------
    Likewise...

    Not too long ago, saw something where someone was (once again) pushing
    for bank-switched registers, and hardware context switching, in the RV
    interrupt handling mechanism.

    One of the benefits of My 66000 having all thread data "effectively"
    reside in memory, is that a simulator can simply read/write simulator
    memory, a small implementation can use a single set of control registers
    and a register file loading and storing like a write back cache, while
    a GBOoO machine can support the model with bank switching. The running software cannot tell the difference.


    OK.

    In the TaskInfo structure, there is a space for saved registers.

    Generally, ISR's like SYSCALL and TLBMISS save the registers there,
    rather than on the ISR stack, mostly because this avoids needing to
    handle them all twice during a context switch.

    Does mean that the TaskInfo needs to be set up before these ISRs can be
    used. Generally, the kernel needs to set up the TaskInfo structure for
    the kernel, then spawn the SYSCALL task, before it can bring up virtual memory.

    This is partly backwards from x86, where one would bring up the MMU first.


    Meanwhile, naive and simple interrupt handling tends to work out
    cheapest and fastest in practice...

    The most important thing to remember here, is to stay as far away
    from x86 as possible. ...


    Errm, I don't think anyone here is proposing bringing back the IDT, GDT,
    and TSS.

    But, yeah, the interrupt mechanism I have is pretty close to the minimum:
    Save off the needed state into a few CRs;
    Set up interrupt mode;
    Which swaps the SP and SSP registers in decoding;
    Generate input entry point address from a base-vector via bit-slicing.


    You don't really want the effective register file to be bigger than it
    needs to be. Likewise, it is not like there is any good way to make a
    register/memory side-channel that is faster than the main register <->
    memory path (via load/store instructions).

    My 66000, due to the way Thread.State is managed, can start loading
    in new state before it stores back current state. Software has to
    store a register before it can load a register. Hardware can load
    a buffer, and then read the old out to another buffer, while writing
    the new in, and then finally migrate the old buffer to memory. Saving
    a round trip latency to wherever the new data is coming from {L2, L3,
    DRAM}.


    OK.

    Probably requires something a little more advanced than the
    direct-mapped cache I am currently using to be effective.

    Though, to make use of loading and saving a context at the same time,
    would mean needing to transfer control without the control first being intercepted by an interrupt handler (meaning the decision to perform a
    context switch would need to itself be driven by hardware).


    Well, unless the ISR is itself another context switch, meaning two such context switches per interrupt...

    At present, the interrupts' context effectively disintegrates/disappears
    when it returns from an interrupt.


    And, if you had an interrupt mechanism that basically just invokes blobs
    of hidden firmware for the interrupt dispatch (and return), this could
    be done, but doesn't gain anything.

    Even I discarded having the control-transfer-switch logic do anything
    other than write-back-cache above. This also means SW can place what
    ever kinds of tables it wants, wherever it wants, of whatever size it
    wants without any page boundaries or alignments getting in the way.

    OK.


    ------------
    So an instruction word starting with 11 but not 1111 is composed of
    11 followed by five six-bit prefix fields.

    Will you object if the compiler writer decides not to use this part of
    the ISA?

    This is where I draw the line, if the compiler can't use it, it does not
    get in--except for reading and writing of control registers.


    Mostly similar.

    I mostly avoid features that would require hand-written ASM or similar
    to use.



    Except for some very niche instructions, like color-cell encoding
    helpers and similar. This being mostly because this is something that
    can happen a lot in some use cases, and needs to be fast.

    In the TestKern GUI, I ended up with a strategy though, of first using a
    quick and dirty ISA driven color-cell encoder, and then if the blocks
    had previously used this one, and haven't changed for a while, it goes
    back over them and re-encodes them with a higher quality but slower SW
    based color-cell encoder.


    Fast strategy:
    Select the min and max colors:
    Internally shuffles the RGB555 bits around and does compares.
    Then, for each block of 4 pixels,
    map its indices between the min and max.

    So, for the RGB555 -> Luma:
    0rrrrrgggggbbbbb => ggrbgrbg (starting at HOBs of each component)
    Generate values, and a comparison matrix, and use this to select.

    To map indices:
    Ymin = rgbtoluma(Cmin);
    Ymax = rgbtoluma(Cmax);
    Ymid = (Ymin + Ymax)>>1;
    Ymlo = (Ymin + Ymid)>>1;
    Ymhi = (Ymid + Ymax)>>1;

    Yc0 = rgbtoluma(Clr0);
    Yc1 = ...

    Ix0 = (Yc0 < Ymid) ?
    ((Yc0 < Ymlo) ? 00 : 01) :
    ((Yc0 < Ymhi) ? 10 : 11) ;
    Ix1 = ...
    Ix2 = ...
    Ix3 = ...

    Then run this logic for each row of pixels.


    In the software encoder, one way to boost quality is to to proper luma
    math, and to select among several possible color axes (typically Cyan/Magenta/Yellow/White), choosing whichever axis has the highest
    contrast.

    Say:
    Ycy = (4*Cb+3*Cg+1*Cr)/8;
    Ymg = (4*Cr+3*Cb+1*Cg)/8;
    Yye = (4*Cg+3*Cr+1*Cb)/8;
    Ywh = (2*cg+1*cr+1*cb)/4;


    Arguably, a general purpose CPU doesn't need this stuff in hardware though.

    So:
    LDTEX
    CCENC
    ...
    Remain as mostly super-niche stuff that only really remains because
    certain use-cases need this stuff to be fast.


    Nevermind whether or not the GUI mode is highly used.


    Decided to leave out a detour about color-cell based video codecs (also
    a bit niche, but for my own uses I was primarily using color-cell based designs rather than MPEG derived designs).


    Well, ironically, apart from my UPIC image format, which ironically is
    an offshoot of my experiments with trying to make MPEG like codecs fast...

    And, UPIC then is basically "What if we took T.81 JPEG and used Rice
    Coding and Block-Haar and the RCT transform...". Ironically, compression
    is still pretty competitive with T.81 JPEG, but is a little faster. It
    also supports a PNG like lossless mode, which ironically for many images
    both beats PNG both on compression and decode speeds.


    The long-standing issue of color-cell designs being basically, how to
    most efficiently encode the color endpoints and block patterns.

    Well, more so in the absence of entropy encoding, because the magic of
    entropy coding coming with the penalty of making everything slow.


    He keeps going at it...

    And we keep humoring him--sad to see is rate of forward progress has
    not changed in 5 years...

    Not really sure the merit of endlessly bashing at non-sane encoding schemes.

    He fails to see he is barking up a large bush instead of a tree.


    Yeah.

    There is a difference between picking a design that basically works.
    And thrashing around with stuff that doesn't.


    In my case, I mostly just seem to be running low on stuff that is
    "actually interesting".


    In my case, I eventually realized I couldn't have *everything* I wanted
    all at the same time, so had dropped off the lower priority items.

    Architecture is as much about what you leave out as what you let in.



    I eventually ended up dropping 16-bit ops, but kept 64 GPRs and predication.

    This was the direction that maximized speed.

    OTOH, 16-bit ops can make sense if the goal is to instead maximizing
    code density.

    But, they are not entirely exclusive:
    XG3 also has OK code density because keeping instruction counts small
    also happens to reduce binary size (even if not the primary goal).

    RV64GC isn't horribly slow, mostly for sake of the 32-bit instructions.


    At one point, XG3 moved into the lead on code density, but RV64GC moved
    back into the lead when I added some experimental 32-bit encodings for:
    LW Xd, Disp16u*4(GP) //256K
    LD Xd, Disp16u*8(GP) //512K

    Both still beating the original form of RV64GC though on code density.


    Though, as some ELF PIE binaries show, there could be a size advantage
    to using a dynamically linked C library, whereas at present
    BGBCC/TestKern is using a static linked C library.


    Otherwise, at least getting closer on trying to get "ld.so" working:
    Sorted out some issues that were breaking the VFS syscalls;
    Also now have mmap seemingly working (*1).

    At present, "ld.so" fails with an assert in its "init_tls" stage.


    *1: It is a little funky ATM as it is needing to pretend to implement a
    4K page size when the underlying CPU isn't actually using 4K pages.
    Seemingly the ELF loader then tries to create multiple overlapping
    mmaps, but the current implementation effectively merges the mmaps
    together in this case.

    Well, because apparently this stuff is partly hard-coded to assume a 4K
    page size and doesn't really work if the page size is different. But,
    does sort of work if the implementation pretends to support 4K pages.

    ...


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Wed Sep 16 05:51:31 2026
    From Newsgroup: comp.arch

    Brian G. Lucas <bagel99@gmail.com> schrieb:
    On 9/14/26 2:16 PM, Thomas Koenig wrote:
    MitchAlsup <user5857@newsgrouper.org.invalid> schrieb:

    BGB <cr88192@gmail.com> posted:

    On 9/12/2026 1:31 PM, John Dallman wrote:
    In article <6aa588d4.1140437@news.eternal-september.org>,
    quadibloc@invalid.com (John Savard) wrote:
    --------------

    Meanwhile... I am still having pretty good results with 64 unified GPRs. >>>>

    The move from XG2 to XG3 had cost a few usable registers:
    R0..R4 have effectively left the GPR space...
    Zero, LR/RA, SP, and GP, TP.

    In My 66000::

    R0 receives the return address on a call, but otherwise is just
    another GPR. In safe-stack mode, R0 is just another GPR.

    It deserves its own register class in a compiler because it acts
    as a substitute for IP as a base register, as "no register" as
    index register, as thread-defined rounding mode in CVTxx, as the
    address to jump back to with "ret" in the absence of a safe stack.

    So it is a bit special.

    The current compiler implementer was lazy and just never allocated R0.

    :-)

    A difference in how LLVM and GCC allocate registers, I assume.

    In GCC, you specify register classes which can be used for certain
    instruction patterns; they also appear as arguments in the __asm
    statements. You would then specify register classes GPRs
    without R0 (for base and index registers, for example) and for
    general usage, for example as operands to ADD. You would also
    specify R0 as the register where subroutine calls drop their
    return address, and that the return address needs to go there.

    You also need tell the GCC the ordering registers are to be
    allocated. It would be natural to put R0 last on that list.

    This could be revisited if register pressure was really an issue.

    My66000 does not need loop unrolling on innermost loops, which is
    good. But Fortran array descriptors, for example, need a lot of
    memory, and unrolling outer loops can also be quite profitable.
    The unified floating point + integer register file could show its
    limits there.

    Right now, we haven't seen that because the test cases were
    plain C, and loop unrolling is something you have to fight
    in LLVM.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2