• Re: adding overflow and carry to every GPR (was Re: OT: Epic RISC-V rant)

    From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 2 17:41:52 2026
    From Newsgroup: comp.arch


    Kragen Javier Sitaker <kragen@canonical.org> posted:

    MitchAlsup <user5857@newsgrouper.org.invalid> writes:
    John Ames <commodorejohn@gmail.com> posted:
    The idea I had last time I dabbled with this stuff was to extend the
    basic add/subtract instructions with an option to save the carry/borrow
    value in another GPR, instead - but I never gave any thought to how to
    represent overflow that way. Hmm.

    That is what CARRY does. The Rd register in CARRY supplies a register
    used to carry the additional values from calculation to calculation.
    The Immediate in CARRY selects which instruction consume Rd {i},
    which instructions generate the next Rd {o} and which do both {io}
    and which do neither {}.

    This is on your 66000?

    Yes, indeed.

    Kragen
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From EricP@ThatWouldBeTelling@thevillage.com to comp.arch on Sat Sep 5 12:08:27 2026
    From Newsgroup: comp.arch

    anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    RISC-V follows its ancestor MIPS (from where it also has many
    mnemonics) in this respect. For a new way to add the carry
    functionality without adding a condition code register, read

    https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf

    (unpublished).

    Looking at example 3 I think the overflow check can be simplified.

    Scheme_Object *ADD_tagged(
    Scheme_Object *tagged_a,
    Scheme_Object *tagged_b)
    {
    intptr_t a = ((intptr_t)tagged_a)>>1;
    intptr_t b = ((intptr_t)tagged_b)>>1;
    intptr_t r;
    Scheme_Object *o;
    r = (uintptr_t)a + (uintptr_t)b;
    o = (Scheme_Object *) ((((uintptr_t)r)<<1)|1);
    r = ((intptr_t )o) >> 1;
    if (b == (uintptr_t)r - (uintptr_t)a)
    return o;
    else
    return ADD_slow (a , b) ;
    }

    tagged_a and tagged_b are converted from 63-bit signed integers
    to 64-bit signed by the arithmetic right shift,
    which also gets rid of the tag lsb.

    The add gives a 64-bit signed result, and you want to check if
    it can be sign-contracted back to 63 bits,
    which is a test if bits r[63] == r[62], which is an XOR.

    intptr_t a = ((intptr_t)tagged_a)>>1;
    intptr_t b = ((intptr_t)tagged_b)>>1;
    intptr_t r, tmp;
    r = a + b;
    tmp = r << 1;
    if ((r ^ tmp) >= 0) // check if bit r[63] == r[62]
    return tmp | 1;
    else
    Add_slow (a, b);





    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.lang.forth,comp.arch on Wed Sep 9 10:50:39 2026
    From Newsgroup: comp.arch

    Kragen Javier Sitaker <kragen@canonical.org> writes: >anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    RISC-V follows its ancestor MIPS (from where it also has many
    mnemonics) in this respect. For a new way to add the carry
    functionality without adding a condition code register, read

    https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf

    (unpublished).

    This is a fascinating idea. I guess the MuP21 had a carry bit in every >"general-purpose register", making them 21 bits,

    <https://www.ultratechnology.com/mup21.html> says this:

    |Internally, MuP21 maintains a 21 bit data/address bus. The MSB bit 20
    |is the carry bit in ALU operations.

    And it mentions an instruction JCZ without explaining it, but given
    that it also has JZ, and the existance of C0 below, I expect that JCZ
    is "jump if carry is zero", while JZ corresponds to T0 (jump if T0-T19
    is false).

    That's pretty much it. I see on
    https://www.ultratechnology.com/f21data.pdf that instruction 03,
    called C0 jumps if T20 (bit 20 of the TOS) is false, and that the
    Forth source would look like "CARRY? IF". F21 is a further
    development of uP21.

    <https://www.ultratechnology.com/mup21.html> has the instruction JCZ

    Thank you for sharing your draft!

    I should finally polish it up with all the parts I left away thanks to
    page limits and publish it on arxiv.

    - anton
    --
    M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
    comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
    New standard: https://forth-standard.org/
    EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.lang.forth,comp.arch on Wed Sep 9 11:16:31 2026
    From Newsgroup: comp.arch

    John Ames <commodorejohn@gmail.com> writes:
    Hm, that *is* a bit interesting, though the idea of including non-
    number info in a numeric register drives me a little crazy.

    The Carry bit is bit 64 of the addition or subtraction of the
    zero-extended operands.

    Concerning the signed oVerflow bit, one could store bit 64 of the
    addition or subtraction of sign-extended operands instead of the
    overflow bit, and then check for a signed overflow by checking if this
    bit 64 is different from bit 63. In that way that bit would also be a
    "number info".

    However, one advantage of storing an overflow bit (i.e., the result of
    this xor) is that we can then combine these bits with "or" or "and".

    The idea I had last time I dabbled with this stuff was to extend the
    basic add/subtract instructions with an option to save the carry/borrow
    value in another GPR, instead

    That has the disadvantage of occupying a write port in the register
    set and in the renamer.

    - but I never gave any thought to how to
    represent overflow that way. Hmm.

    A GPR has many bits:-)

    - anton
    --
    M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
    comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
    New standard: https://forth-standard.org/
    EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Wed Sep 9 13:36:51 2026
    From Newsgroup: comp.arch

    EricP <ThatWouldBeTelling@thevillage.com> writes:
    anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf

    (unpublished).

    Looking at example 3 I think the overflow check can be simplified.

    Scheme_Object *ADD_tagged(
    Scheme_Object *tagged_a,
    Scheme_Object *tagged_b)
    {
    intptr_t a = ((intptr_t)tagged_a)>>1;
    intptr_t b = ((intptr_t)tagged_b)>>1;
    intptr_t r;
    Scheme_Object *o;
    r = (uintptr_t)a + (uintptr_t)b;
    o = (Scheme_Object *) ((((uintptr_t)r)<<1)|1);
    r = ((intptr_t )o) >> 1;
    if (b == (uintptr_t)r - (uintptr_t)a)
    return o;
    else
    return ADD_slow (a , b) ;
    }

    This code in the upper part of Figure 3 is not from me, but from the
    Racket 8.6 source code, and footnote 5 explains the following:

    |The shown code is derived from the original ADD function by putting
    |the untagging at the callee rather than the caller side, expanding the |macros, and simplifying the resulting code. The C and assembler code
    |for the lower part can also be found at
    |https://godbolt.org/z/TroqhsrrM

    So one can imagine why the Racket 8.6 implementors have not made the optimization that you suggest, nor the optimization that I suggest.
    Also, for Racket 8.6 this was only part of the legacy BC (bytecode) implementation, with their native-code compiler being where the action
    was, but I guess that this missed optimization existed before the
    native-code compiler was added (and who knows how this case is handled
    there; I did not find it).

    tagged_a and tagged_b are converted from 63-bit signed integers
    to 64-bit signed by the arithmetic right shift,
    which also gets rid of the tag lsb.

    The add gives a 64-bit signed result, and you want to check if
    it can be sign-contracted back to 63 bits,
    which is a test if bits r[63] == r[62], which is an XOR.

    intptr_t a = ((intptr_t)tagged_a)>>1;
    intptr_t b = ((intptr_t)tagged_b)>>1;
    intptr_t r, tmp;
    r = a + b;
    tmp = r << 1;
    if ((r ^ tmp) >= 0) // check if bit r[63] == r[62]
    return tmp | 1;
    else
    Add_slow (a, b);

    Yes, I thought about this optimization, too, but also came up with:

    intptr_t a1=(intptr_t) tagged_a;
    intptr_t b1=(intptr_t) tagged_b;
    intptr_t r ;

    if (!__builtin_add_overflow(a1,(b1-1),&r))
    return (Scheme_Object *)r;

    intptr_t a = ((intptr_t)tagged_a)>>1;
    intptr_t b = ((intptr_t)tagged_b)>>1;
    return ADD_slow (a , b );

    which you see in the lower part of Figure 3.

    On AMD64 this compiles to:

    leaq -1(%rsi), %rax
    addq %rdi, %rax
    jo .L8
    ret
    .L8:
    sarq %rsi
    sarq %rdi
    jmp ADD_slow

    I.e., just 4 instructions in the common case (reducible to 3 by
    choosing 0 as the tag for small integers).

    It's interesting to think about what it would take to enable an
    optimizer to optimize the former, or the original Racket code, into
    the assembly code above, or what would come out of your code. I guess
    that if the BC run-time of Racket were part of the next SPEC CPU
    suite, and a Scheme program with lots of small additions in the
    reference input, we would see:-).

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Rubin@no.email@nospam.invalid to comp.lang.forth,comp.arch on Wed Sep 9 22:10:37 2026
    From Newsgroup: comp.arch

    anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    I should finally polish it up with all the parts I left away thanks to
    page limits and publish it on arxiv.

    I remember we had a discussion of this idea here on clf, some of which
    spilled to comp.arch. I don't know if it contained anything that would
    be useful in your paper that's not already in it.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Kragen Javier Sitaker@kragen@canonical.org to comp.lang.forth,comp.arch on Fri Sep 11 13:48:30 2026
    From Newsgroup: comp.arch

    John Ames <commodorejohn@gmail.com> writes:
    On Wed, 02 Sep 2026 02:04:51 GMT
    MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
    For ADD, carry Rd contains only 1 bit, but for others, it contains
    rCLall the bits that did not make it into the standard resultrCY. This
    makes CARRY suitable for all integer multiprecision arithmetic and
    for significant quantities of FP arithmetic.
    [...]
    Multi-precisions shifts use carry

    Yes, exactly - add/subtract set the designated overflow/carry register
    to 1 on a carry/borrow (which can then be added/subtracted from the next
    more significant word) but e.g. multiply can set it to the whole most- significant half of the product, multi-place shifts can catch all the
    bits that fall off one end in the other so that rotate becomes shift-
    and-OR, etc. Very flexible, especially if you have the capability to
    branch on the value of a GPR.

    A bit-shift operation with two source registers is sometimes called a
    rCLfunnel shiftrCY, which is sort of the inverse of this: <https://www.eevblog.com/forum/programming/funnel-shifter-where-are-useful/> <https://stackoverflow.com/questions/12767113/funnel-shift-what-is-it> <https://github.com/rust-lang/rust/issues/145686>

    That is, instead of one source register (for a given bit shift size) and
    two destination registers, a funnel shift has two source registers (plus
    the amount to shift) and possibly just one destination register.

    An interesting thing about funnel shifts is that, if both source
    registers are the same, it provides bit rotation, not just shift.
    Conceivably you could get the same result from this inverse funnel
    shift, depending on how you handle the case where the destination
    register is the same as the CARRY register.

    Kragen
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Kragen Javier Sitaker@kragen@canonical.org to comp.lang.forth,comp.arch on Fri Sep 11 14:28:06 2026
    From Newsgroup: comp.arch

    anton@mips.complang.tuwien.ac.at (Anton Ertl) writes:
    Kragen Javier Sitaker <kragen@canonical.org> writes:
    Thank you for sharing your draft!

    I should finally polish it up with all the parts I left away thanks to
    page limits and publish it on arxiv.

    I agree wholeheartedly.

    Kragen
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.lang.forth,comp.arch on Fri Sep 11 18:03:54 2026
    From Newsgroup: comp.arch


    Kragen Javier Sitaker <kragen@canonical.org> posted:

    John Ames <commodorejohn@gmail.com> writes:
    On Wed, 02 Sep 2026 02:04:51 GMT
    MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
    For ADD, carry Rd contains only 1 bit, but for others, it contains
    rCLall the bits that did not make it into the standard resultrCY. This
    makes CARRY suitable for all integer multiprecision arithmetic and
    for significant quantities of FP arithmetic.
    [...]
    Multi-precisions shifts use carry

    Yes, exactly - add/subtract set the designated overflow/carry register
    to 1 on a carry/borrow (which can then be added/subtracted from the next more significant word) but e.g. multiply can set it to the whole most- significant half of the product, multi-place shifts can catch all the
    bits that fall off one end in the other so that rotate becomes shift- and-OR, etc. Very flexible, especially if you have the capability to
    branch on the value of a GPR.

    A bit-shift operation with two source registers is sometimes called a rCLfunnel shiftrCY, which is sort of the inverse of this: <https://www.eevblog.com/forum/programming/funnel-shifter-where-are-useful/> <https://stackoverflow.com/questions/12767113/funnel-shift-what-is-it> <https://github.com/rust-lang/rust/issues/145686>

    That is, instead of one source register (for a given bit shift size) and
    two destination registers, a funnel shift has two source registers (plus
    the amount to shift) and possibly just one destination register.

    An interesting thing about funnel shifts is that, if both source
    registers are the same, it provides bit rotation, not just shift.
    Conceivably you could get the same result from this inverse funnel
    shift, depending on how you handle the case where the destination
    register is the same as the CARRY register.

    Yes, CARRY is how one gets shift double-wide in My 66000. So, when we
    added Rotates, we use a length field (rather than a 2^(3+n) specifier)
    so, one can rotate a 6-bit field or a 19-bit field in either direction.

    Kragen
    --- Synchronet 3.22a-Linux NewsLink 1.2