• interrupt handling latency and weird register use (was Re: OT: Epic RISC-V rant)

    From Kragen Javier Sitaker@kragen@canonical.org to comp.lang.forth,comp.arch,alt.lang.asm on Tue Sep 1 17:00:23 2026
    From Newsgroup: alt.lang.asm

    antispam@fricas.org (Waldek Hebisch) writes:
    Paul Rubin <no.email@nospam.invalid> wrote:
    https://dmitry.gr/?r=06.%20Thoughts&proj=12.%20RV

    Some comments.

    1) Interrupt latency claim is half-truth. First, normal RISC-V has
    32 registers while ARM has 16. If you can do with 16 registers
    divide register set into 2 parts, use one part for normal code
    and the other for interrupt handler. That way you will get few
    cycle overhead, impossible feat for Cortex M. So it is really
    a tradeoff; do you want to have modest overhead for pretty
    typical use case, or do you for very low overhead in cases that
    need it and higher overhead for typical cases?

    This is a really good point, and one I should have thought of. You can
    reserve some registers as rCLFIQrCY registers and only use them inside interrupt handlers, depsite the absence of an architectural FIQ
    mechanism (which was present on the ARM2 but excised in the Cortex-M).

    (Note that most existing RISC-V cores like the QingKe core used in the
    CH32V003 have their own, incompatible, interrupt-handling mechanisms.
    Generally these are FIQ-like, enabling lower latency than the standard
    RISC-V mechanism but with a limited depth of nested interrupts.)

    You can write the rCLFIQrCY handler itself in assembly, since if it needs to
    be more than about 16 instructions long you might as well switch to the standard ABI, but you also need to compile the rest of your application
    with the rCLFIQ registersrCY reserved, including any system libraries you
    might be using, such as an integer division subroutine. This is true
    whether yourCOre using Forth, or C, or any other language.

    8 registers is probably enough for the FIQ handler, so with RV32I you
    have 24 left over for the rest of the code.

    Reserving some registers is a very small modification to most compilers
    (those that do some kind of register allocation), and Forth compilers
    are fairly small, but making any modification to GCC or LLVM is a bit
    daunting, even writing a new rCLmachine definitionrCY file. I had
    previously looked for an existing way to do this with GCC and given up,
    but it turns out that itrCOs fairly simple; GCC has an `-ffixed-reg`
    option, where `reg` is the name of the register to reserve. See <https://gcc.gnu.org/onlinedocs/gcc/Code-Gen-Options.html>.

    So, in theory, you ought to be able to reserve x24 through x31 for rCLFIQrCY handlers by compiling all your C source code with something like

    gcc -ffixed-x24 -ffixed-x25 -ffixed-x26 -ffixed-x27 -ffixed-x28 \
    -ffixed-x29 -ffixed-x30 -ffixed-x31

    I have verified that GCC does indeed accept these command-line options,
    and that it respects -ffixed-a5, but I havenrCOt written code with
    sufficient register pressure to provoke GCC to try to use x24 (s8) even
    without this option. So, itrCOs documented to work, but itrCOs kind of a
    niche feature, and I havenrCOt tried it in practice.

    You might be tempted to try to use GCCrCOs global register variables <https://gcc.gnu.org/onlinedocs/gcc/Global-Register-Variables.html> for
    this:

    register int *foo asm ("r12");

    However, the documentation specifically says that you should use
    `-ffixed-reg` instead, because global register variables will not work
    for the interrupt-handling case:

    Similarly, it is not safe to access the global register variables
    from signal handlers or from more than one thread of control. Unless
    you recompile them specially for the task at hand, the system
    library routines may temporarily use the register for other things.
    Furthermore, since the register is not reserved exclusively for the
    variable, accessing it from handlers of asynchronous signals may
    observe unrelated temporary values residing in the register.

    For this particular case, another possibility may be to compile for
    RV32E, which only has 16 architectural general-purpose registers but is otherwise compatible with RV32I.

    I had looked for how to achieve this three years ago for a project
    called Monokokko, which is a five-machine-instruction-long cooperative-multitasking OS for ARM:

    .thumb_func
    yield: push {r4-r9, r11, lr} @ save all callee-saved regs except r10
    str sp, [r10], #4 @ save stack pointer in current task
    ldr r10, [r10] @ load pointer to next task
    ldr sp, [r10] @ switch to next task's stack
    pop {r4-r9, r11, pc} @ return into yielded context there

    <http://canonical.org/~kragen/sw/dev3/monokokko.S>

    Now I know how to solve the problem! So I can write Monokokko tasks in
    C now!

    5) Optionality. I do not like it and it is probably biggest
    problem of RISC-V. OTOH to have any chance of success RISC-V
    need a buy-in from several independent parties. I suspect that
    the only practical way to get consensus from varied parties is
    by making most of specification optional.

    For CPU vendors and computer architecture researchers, I suspect this is
    the biggest selling point of RISC-V: they can add experimental vector extensions without waiting for them to be ratified, they can add
    FIQ-style low-latency interrupt handling, they can implement their own
    memory protection mechanisms, they can add zero-overhead loops, etc. As
    I understand it, Berkeley researchersrCO nightmarish negotiations with ARM
    to get a license to perform such experiments on ARM cores was the hell
    from which RISC-V came in the first place.

    It also means that bad decisions by RISC-V International have much less
    impact on licensees. The most obvious example, to my mind, is something GrinbergrCOs critique doesnrCOt even mentionrCerCorCeitrCOs having the divide instruction in the M extension. Hardware multiplication is crucial for
    all kinds of applications, including software-defined radio and other
    DSP, image processing, 3-D rendering, and neural networks. By contrast, hardware division is a minor advantage but rarely appears in inner
    loops, and has been omitted from historical architectures including the
    Cray-1, Cray-2, and ARM.

    Kragen
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.lang.forth,comp.arch,alt.lang.asm on Tue Sep 1 20:51:46 2026
    From Newsgroup: alt.lang.asm

    Kragen Javier Sitaker <kragen@canonical.org> writes:
    The most obvious example, to my mind, is something
    GrinbergrCOs critique doesnrCOt even mentionrCerCorCeitrCOs having the divide >instruction in the M extension. Hardware multiplication is crucial for
    all kinds of applications, including software-defined radio and other
    DSP, image processing, 3-D rendering, and neural networks. By contrast, >hardware division is a minor advantage but rarely appears in inner
    loops, and has been omitted from historical architectures including the >Cray-1, Cray-2, and ARM.

    And Alpha and, most recently, IA-64.

    Grinberg actually mentioned it, but got it wrong. He claimed that multiplication and division are indepenently optional, while in the
    RISC-V specification, you get none or both.

    I found that surprising, too, but the world has continued to turn
    since the days of Alpha and IA-64, so I guess they thought: Any CPU
    big enough to have a multiplier will also be big enough for a division instruction.

    Followups set to comp.arch.

    - anton
    --
    M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
    comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
    New standard: https://forth-standard.org/
    EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.lang.forth,comp.arch,alt.lang.asm on Wed Sep 2 01:43:26 2026
    From Newsgroup: alt.lang.asm


    Kragen Javier Sitaker <kragen@canonical.org> posted:

    antispam@fricas.org (Waldek Hebisch) writes:
    Paul Rubin <no.email@nospam.invalid> wrote:
    https://dmitry.gr/?r=06.%20Thoughts&proj=12.%20RV

    Some comments.

    1) Interrupt latency claim is half-truth. First, normal RISC-V has
    32 registers while ARM has 16. If you can do with 16 registers
    divide register set into 2 parts, use one part for normal code
    and the other for interrupt handler.-------snip-------

    This is a really good point, and one I should have thought of. You can reserve some registers as rCLFIQrCY registers and only use them inside interrupt handlers, depsite the absence of an architectural FIQ
    mechanism (which was present on the ARM2 but excised in the Cortex-M).

    My 66000 has HW perform the register movements (save and restore) as
    if the RF was a write back cache. This enables HW to start saving
    registers before the first instruction in ISR has been fetched
    (which, necessarily, is after the address of the ISR has been loaded).
    For interrupts all this pre-loading can transpire in parallel with
    the storing of the registers to be pushed out.

    For SVCs, R0-R8 are arguments to the SVC, while R9-R15 are not saved
    {just like R9-R16 is not preserved in a normal procedure call}. An
    SVR <return> R1-R2 contain returning arguments while R3-R15 are cleared.
    So, supervisor can see noise in R3-R15 but caller cannot see noise after return. R16-R31 are preserved across the SVC-SVR just like between CALL
    and RET.

    It also has the benefit that once control arrives in ISR, you have
    <at least> 16 registers with useful state {SP to ISR stack, FP,
    and various pointers to structures the ISR might need} without
    having to load them with instructions.

    While "IN" an ISR, the ISR might need to schedule a DPC or softIRQ.
    My 66000 allows this to become manifest with a single instruction
    (a store to the Interrupt aperture with a message describing where
    the DPC/softIRQ arguments are to be found.) And a few instructions
    setting up the deferred arguments.

    (Note that most existing RISC-V cores like the QingKe core used in the CH32V003 have their own, incompatible, interrupt-handling mechanisms.

    Almost always a bad decision.

    Generally these are FIQ-like, enabling lower latency than the standard
    RISC-V mechanism but with a limited depth of nested interrupts.)

    You can write the rCLFIQrCY handler itself in assembly, since if it needs to be more than about 16 instructions long you might as well switch to the standard ABI, but you also need to compile the rest of your application
    with the rCLFIQ registersrCY reserved, including any system libraries you might be using, such as an integer division subroutine. This is true
    whether yourCOre using Forth, or C, or any other language.

    In My 66000, There is a 6 instruction dispatcher (could be per Thread
    or common across all threads of a given privilege) that is written in
    ASM, but the ISR is written in any suitable HLL.

    I went with a dispatcher model so that SW gets to decide where the
    tables are, how they are configured, sized, and accessed--without
    arbitrary boundaries (page alignment page size).

    8 registers is probably enough for the FIQ handler, so with RV32I you
    have 24 left over for the rest of the code.

    A 24 register machine will perform ~6% worse than a 32-real register
    machine. Can you make up in responsiveness of interrupts for this
    loss ?? Most RISCs are not 32-real register machines. R0 containing
    the value 0, certain registers dedicated to dynamic linking (MIPS),
    R1 containing the return address, and any special registers that have
    to do with multiply and divide (sometimes double wide shifts).

    ----------
    I had looked for how to achieve this three years ago for a project
    called Monokokko, which is a five-machine-instruction-long cooperative-multitasking OS for ARM:

    .thumb_func
    yield: push {r4-r9, r11, lr} @ save all callee-saved regs except r10
    str sp, [r10], #4 @ save stack pointer in current task
    ldr r10, [r10] @ load pointer to next task
    ldr sp, [r10] @ switch to next task's stack
    pop {r4-r9, r11, pc} @ return into yielded context there

    <http://canonical.org/~kragen/sw/dev3/monokokko.S>

    // you want to pass control to thread in R19 ADA-like
    CR R19,Application,R19
    // R19 now points to yourself, but control is now there

    Now I know how to solve the problem! So I can write Monokokko tasks in
    C now!

    5) Optionality. I do not like it and it is probably biggest
    problem of RISC-V. OTOH to have any chance of success RISC-V
    need a buy-in from several independent parties. I suspect that
    the only practical way to get consensus from varied parties is
    by making most of specification optional.

    For CPU vendors and computer architecture researchers, I suspect this is
    the biggest selling point of RISC-V: they can add experimental vector extensions without waiting for them to be ratified, they can add
    FIQ-style low-latency interrupt handling, they can implement their own
    memory protection mechanisms, they can add zero-overhead loops, etc. As
    I understand it, Berkeley researchersrCO nightmarish negotiations with ARM
    to get a license to perform such experiments on ARM cores was the hell
    from which RISC-V came in the first place.

    And with that attitude ARM deserves the competition.

    It also means that bad decisions by RISC-V International have much less impact on licensees. The most obvious example, to my mind, is something GrinbergrCOs critique doesnrCOt even mentionrCerCorCeitrCOs having the divide instruction in the M extension. Hardware multiplication is crucial for
    all kinds of applications, including software-defined radio and other
    DSP, image processing, 3-D rendering, and neural networks. By contrast, hardware division is a minor advantage but rarely appears in inner
    loops, and has been omitted from historical architectures including the Cray-1, Cray-2, and ARM.

    This leaves out the CRAY solution of a reciprocation instruction making
    FDIV to be RCP->FMUL. {{Handwaving about how you cannot achieve IEEE 754 precision and accuracy doing it that way...}}

    FDIV is so bad that many GPU languages carry around w=1/SQRT(x^2+y^2+z^2)
    so that for all intents and purposes, FDIV is a FMUL.

    But all languages support FDIV, thereby, FDIV should be an available instruction--differing implementations can provide for this using any
    means chosen by implementation; including {trap to SW, Slow FU, medium
    FU, fast FU, SIMD FU, ...} {{{If you choose 'trap to SW'--you better
    haven a clean trapping mechanism}}}

    Kragen
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Andy Valencia@vandys@vsta.org to comp.lang.forth,comp.arch,alt.lang.asm on Wed Sep 2 08:11:15 2026
    From Newsgroup: alt.lang.asm

    Kragen Javier Sitaker <kragen@canonical.org> writes:
    ... You can
    reserve some registers as "FIQ" registers and only use them inside
    interrupt handlers, depsite the absence of an architectural FIQ
    mechanism (which was present on the ARM2 but excised in the Cortex-M).

    This brings back memories of MIPS' K0 and K1 registers. From non-interrupt code perspective their value could change asynchronously, as they were
    there to be scratched upon by the lowest level of interrupt and trap
    entry. It always struck me as a covert channel, though I never came up
    with a way to use it.

    Andy Valencia
    Home page: https://www.vsta.org/andy/
    To contact me: https://www.vsta.org/contact/andy.html
    No AI was used in the composition of this message
    --- Synchronet 3.22a-Linux NewsLink 1.2