• A New Kind of Insanity

    From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Wed Sep 23 00:53:14 2026
    From Newsgroup: comp.arch

    I tried to find a way to fit everything I wanted into too little
    opcode space. But all the strategems I tried were, rightly, criticized
    as being too complex, as being something people would never accept.

    So what else is left? In the past, I had tried various ways of shaving
    one bit off the length of instructions through various compromises. I
    was dissatisfied with them all, and went around in circles for so long
    with that.

    So I thought I would never go back there. But it seems there's no
    other choice.

    Thus, I am now going to get close to having everything I wanted - but
    not closer than is possible. The scheme is now:

    0xx where xx is not 11: 32-bit instructions. With 16-bit
    displacements, the base register field will only be two bits long,
    allowing three base registers with these instructions.

    1: 16-bit instructions. The source and destination registers can only
    be registers 0 to 15. This lets them fit.

    0110 - 48-bit instructions.
    01110 - 64-bit instructions.

    Among the 64-bit instructions, there will be an instruction perhaps of
    the form 0111011 followed by three 19-bit slots. These slots might
    contain short instructions that now can work with all 32 registers for
    both operands without restriction, and even set the condition codes.

    Or two of those slots could contain a slightly larger than 32-bit
    instruction - including memory-reference instructions with all seven
    base registers for 32-bit displacements, and memory-to-register
    operate instructions.

    But because those 64-bit instructions can start at any 16-bit
    position, there's a problem. This new "complete" instruction set is
    crippled if it is impossible to branch into a sequence of such
    instructions except at every third position.
    So I need a special branch instruction with an extra two bits. It
    probably won't fit in the regular 32-bit instructions, but in the
    38-bit instructions, that would be fine.

    Oops. One big problem. So the fancy jump instruction has, as its
    effective address, the start of the 64-bit instruction containing
    19-bit units. Then there's an extra two bits to pick which one. (Only
    one bit is "really" needed, since for the first unit, you can just
    branch to the 64-bit instruction. Wasting a bit for ease of
    comprehension, though, seems appropriate.)

    What's the big problem? When do you have *only* an address to
    determine where you're branching to, and can't have other fields in
    the instruction involved specifying anything else?
    When you're making use of the return address given to a subroutine
    upon entry!

    So here is the insanity I still must engage in.
    Return addresses must include special coding. If the return location
    is to an instruction that begins within one of these 64-bit capsule instructions... its form must be as follows:
    (2 bits): Either 10 or 11, for the second or third 19-bit slot.
    Possibly 00 and 01, which are synonymous, for the first one, in the
    case of some freak interrupt.
    (61 bits): The actual bits specifying the address of the 64-bit
    instruction containing those slots.
    1: All instructions are multiples of 16 bits in length, and they all
    begin on 16-bit boundaries. So odd addresses don't exist; this bit is, therefore, available as the flag to indicate a "special" address that
    both points to a 64-bit instruction and to a slot within it.

    I think that it's possible to get by with merely 62-bit addressing.
    Of course, though, this means that in 32-bit mode, suddenly the
    address space is limited to *one* measly gigabyte instead of four! So
    in the 32-bit days, this would have been painful indeed.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Wed Sep 23 18:14:22 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:

    I tried to find a way to fit everything I wanted into too little
    opcode space. But all the strategems I tried were, rightly, criticized
    as being too complex, as being something people would never accept.

    So what else is left? In the past, I had tried various ways of shaving
    one bit off the length of instructions through various compromises. I
    was dissatisfied with them all, and went around in circles for so long
    with that.

    More than 4 years and counting...or is it 6? In any event convergence is
    slow.

    So I thought I would never go back there. But it seems there's no
    other choice.

    Thus, I am now going to get close to having everything I wanted - but
    not closer than is possible. The scheme is now:

    0xx where xx is not 11: 32-bit instructions. With 16-bit
    displacements, the base register field will only be two bits long,
    allowing three base registers with these instructions.

    1: 16-bit instructions. The source and destination registers can only
    be registers 0 to 15. This lets them fit.

    0110 - 48-bit instructions.
    01110 - 64-bit instructions.

    I think part of the problem is that you continue to lie to yourself.

    111yyyyyyyyyyyyy is a 13-bit instruction with a 3-bit marker. 1xxyyyyyyyyyyyyyyyyyyyyyyyyyyyyy is a 29-bit instruction with a 3-bit
    marker.
    0110 is a 44-bit instruction with a 4-bit marker.
    01110 is a 59-bit instruction with a 5-bit marker.

    Among the 64-bit instructions, there will be an instruction perhaps of
    the form 0111011 followed by three 19-bit slots. These slots might
    contain short instructions that now can work with all 32 registers for
    both operands without restriction, and even set the condition codes.

    Or two of those slots could contain a slightly larger than 32-bit
    instruction - including memory-reference instructions with all seven
    base registers for 32-bit displacements, and memory-to-register
    operate instructions.

    But because those 64-bit instructions can start at any 16-bit
    position, there's a problem. This new "complete" instruction set is
    crippled if it is impossible to branch into a sequence of such
    instructions except at every third position.

    Bingo !

    So I need a special branch instruction with an extra two bits. It
    probably won't fit in the regular 32-bit instructions, but in the
    38-bit instructions, that would be fine.

    Or avoid the mess.

    Oops. One big problem. So the fancy jump instruction has, as its
    effective address, the start of the 64-bit instruction containing
    19-bit units. Then there's an extra two bits to pick which one. (Only
    one bit is "really" needed, since for the first unit, you can just
    branch to the 64-bit instruction. Wasting a bit for ease of
    comprehension, though, seems appropriate.)

    What's the big problem? When do you have *only* an address to
    determine where you're branching to, and can't have other fields in
    the instruction involved specifying anything else?
    When you're making use of the return address given to a subroutine
    upon entry!

    Do not have a CALL instruction that is not on an appropriate boundary,
    so, the subsequent instruction is not also on an appropriate boundary.

    So here is the insanity I still must engage in.

    Remember:: Nobody else did this to you.

    Return addresses must include special coding. If the return location
    is to an instruction that begins within one of these 64-bit capsule instructions... its form must be as follows:
    (2 bits): Either 10 or 11, for the second or third 19-bit slot.
    Possibly 00 and 01, which are synonymous, for the first one, in the
    case of some freak interrupt.
    (61 bits): The actual bits specifying the address of the 64-bit
    instruction containing those slots.
    1: All instructions are multiples of 16 bits in length, and they all
    begin on 16-bit boundaries. So odd addresses don't exist; this bit is, therefore, available as the flag to indicate a "special" address that
    both points to a 64-bit instruction and to a slot within it.

    I think that it's possible to get by with merely 62-bit addressing.

    Likely...

    Of course, though, this means that in 32-bit mode, suddenly the
    address space is limited to *one* measly gigabyte instead of four! So
    in the 32-bit days, this would have been painful indeed.

    Why would an architecture, designed in 2026, have a 32-bit address "mode".

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Thu Sep 24 04:48:02 2026
    From Newsgroup: comp.arch

    On Wed, 23 Sep 2026 18:14:22 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    I think part of the problem is that you continue to lie to yourself.

    111yyyyyyyyyyyyy is a 13-bit instruction with a 3-bit marker. >1xxyyyyyyyyyyyyyyyyyyyyyyyyyyyyy is a 29-bit instruction with a 3-bit
    marker.
    0110 is a 44-bit instruction with a 4-bit marker.
    01110 is a 59-bit instruction with a 5-bit marker.

    The instructions take up a certain amount of space, and they have a
    certain number of free useful bits in them. Both facts are important.

    Do not have a CALL instruction that is not on an appropriate boundary,
    so, the subsequent instruction is not also on an appropriate boundary.

    I think that I have a good answer to _this_ objection, even if not to
    any of the others.

    This kind of restriction - not being able to place instructions in
    arbitrary positions - is _exactly_ what people, including, I think,
    yourself, have complained about before. If a capability requires that
    sort of contrivance to use it, then compilers will never use it.

    Why would an architecture, designed in 2026, have a 32-bit address "mode".

    This one, I have to admit, I don't really have a good answer for,
    since I haven't bothered to define, though I could have, 32-bit
    versions of the last few Concertina iterations.
    Some machines do have less than a gigabyte of memory, and some
    programs are smaller than a gigabyte in size. So if such programs used
    32-bit addressing, they would save memory, and save on time and/or
    power consumption when doing address arithmetic.
    And there are existing architectures with 32-bit address modes as
    legacy cruft if nothing else, with which I would have to compete, and
    so _if_ the lack of it were a problem, _then_ perhaps I should avoid
    it.
    But having multiple instruction modes is a security risk, and the
    savings from 32-bit mode are minimal, so I tend to agree that I don't
    really need to worry about the fact that my subroutine return scheme
    has an impact that _would_ have been significant in a 32-bit world.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Thu Sep 24 05:03:19 2026
    From Newsgroup: comp.arch

    On Wed, 23 Sep 2026 18:14:22 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    Remember:: Nobody else did this to you.

    But am I doing it to myself out of sheer masochism?

    No: instead, I'm paying the price for my greed, which is a different
    situation.

    The problem is: I'd like to have an architecture in which the 32-bit instructions do "this", and the 16-bit instructions do "that", with my standards for this and that being based on the IBM System/360.

    But the IBM System/360 had 16 general registers (and four
    floating-point registers, but in ESA/390 it was easy to expand that to
    16 as well). I want to have 32 registers in the integer and floating
    register banks because RISC chips that are popular these days have
    that.

    So I am engaging in insane schemes basically for the purpose of
    cheating the Pigeonhole Principle.

    It would be simpler and more orthogonal if I were to propose an
    architecture which used bit addressing, like the IBM 7030/STRETCH, but
    in which, unlike the STRETCH, _instructions_ were arbitrary numbers
    of bits in length.

    So if memory-reference instructions that can do what I want happen to
    take up 37 bits, no problem. They can be 37 bits long.

    I presume, though, that such a level of barrel-shifter overuse would
    be inefficient. So I'm instead trying to come up with schemes that
    play nicely with power-of-two buses and memory.

    Thus, I implement instructions that are slightly bigger than 16 bits
    and instructions that are slightly bigger than 32 bits... by packing
    three 18-bit instruction chunks inside a 64-bit instruction.
    The rest is all a consequence of respecting the advice that said that
    building an ISA around 256-bit instruction blocks is bad, and
    variable-length instructions don't need to be avoided for efficiency.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 24 17:18:03 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:

    On Wed, 23 Sep 2026 18:14:22 GMT, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    Remember:: Nobody else did this to you.

    But am I doing it to myself out of sheer masochism?

    No: instead, I'm paying the price for my greed, which is a different situation.

    But; you ARE doing it to yourself.

    The problem is: I'd like to have an architecture in which the 32-bit instructions do "this", and the 16-bit instructions do "that", with my standards for this and that being based on the IBM System/360.

    But the IBM System/360 had 16 general registers (and four
    floating-point registers, but in ESA/390 it was easy to expand that to
    16 as well). I want to have 32 registers in the integer and floating
    register banks because RISC chips that are popular these days have
    that.

    So I am engaging in insane schemes basically for the purpose of
    cheating the Pigeonhole Principle.

    You are also letting perfection get in the way of good enough.

    It would be simpler and more orthogonal if I were to propose an
    architecture which used bit addressing, like the IBM 7030/STRETCH, but
    in which, unlike the STRETCH, _instructions_ were arbitrary numbers
    of bits in length.

    Somewhat Mill-like.

    So if memory-reference instructions that can do what I want happen to
    take up 37 bits, no problem. They can be 37 bits long.

    I presume, though, that such a level of barrel-shifter overuse would
    be inefficient. So I'm instead trying to come up with schemes that
    play nicely with power-of-two buses and memory.

    Thus, I implement instructions that are slightly bigger than 16 bits
    and instructions that are slightly bigger than 32 bits... by packing
    three 18-bit instruction chunks inside a 64-bit instruction.

    If you tried really hard, you could get 3|u20-bit instructions in 64-
    bits, and 2|u15-bits in 32-bits.

    My suggestion is fewer container sizer = easier Decoding and easier
    boundary controls.

    The rest is all a consequence of respecting the advice that said that building an ISA around 256-bit instruction blocks is bad, and
    variable-length instructions don't need to be avoided for efficiency.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Thu Sep 24 19:16:40 2026
    From Newsgroup: comp.arch

    On Thu, 24 Sep 2026 17:18:03 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    My suggestion is fewer container sizer = easier Decoding and easier
    boundary controls.

    So far, I only have one container size.
    While a previous proposal did have three 20-bit instructions in 64
    bits, now I have to get by with fewer, since I have other instructions
    for which to make room.
    Seven bits for opcode, one bit to indicate if the condition codes are
    set, ten bits for source and destination registers - 18 bits is enough
    for a full short instruction. So reserving three 19-bit instruction
    slots in 64 bits lets me also have room for 36-bit instructions.
    Oh, dear. One bit isn't enough to distinguish between: 18-bit short,
    start of 36-bit, and do not decode.
    20 bits is possible, I just have to assign fully half of the opcode
    space starting with 011 to these containers. Which is fine; they're
    important, for making the instruction set extensible.
    The container mechanism does allow, since it branches to the container instruction, then indicating an instruction within it, for other
    container types. I am envisioning _one_ other container type, with
    longer instruction slots, to handle instructions with an explicit
    indication of parallelism, and predication.

    The idea is to have an ISA that is almost indefinitely extensible - at
    the cost of relatively high levels of overhead for the extensions.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Thu Sep 24 22:06:51 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:

    On Thu, 24 Sep 2026 17:18:03 GMT, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    My suggestion is fewer container sizer = easier Decoding and easier >boundary controls.

    So far, I only have one container size.

    You stated 5 different instruction container sizes packed into one fetch-container.

    While a previous proposal did have three 20-bit instructions in 64
    bits, now I have to get by with fewer, since I have other instructions
    for which to make room.

    ... by what is left out ...

    Seven bits for opcode, one bit to indicate if the condition codes are
    set, ten bits for source and destination registers - 18 bits is enough
    for a full short instruction. So reserving three 19-bit instruction
    slots in 64 bits lets me also have room for 36-bit instructions.
    Oh, dear. One bit isn't enough to distinguish between: 18-bit short,
    start of 36-bit, and do not decode.

    It is just not getting through is it ???

    20 bits is possible, I just have to assign fully half of the opcode
    space starting with 011 to these containers. Which is fine; they're important, for making the instruction set extensible.
    The container mechanism does allow, since it branches to the container instruction, then indicating an instruction within it, for other
    container types. I am envisioning _one_ other container type, with
    longer instruction slots, to handle instructions with an explicit
    indication of parallelism, and predication.

    The idea is to have an ISA that is almost indefinitely extensible - at
    the cost of relatively high levels of overhead for the extensions.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Thu Sep 24 22:29:02 2026
    From Newsgroup: comp.arch

    On Thu, 24 Sep 2026 22:06:51 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    You stated 5 different instruction container sizes packed into one >fetch-container.

    I'm afraid I don't understand what you mean here. Are you referring to
    an older idea from another thread?

    Normal code is CISC, The first few bits of an instruction indicate the
    length of an instruction.

    0xx, where xx are not both 1, indicates a 32-bit instruction.
    1 indicates a 16-bit instruction.
    0111 will now have to indicate a 64-bit container, since it seems like
    I will require three 20-bit slots in each container, and 19-bit slots,
    for example, will not be enough.
    So now
    01100 will begin a 48-bit instruction,
    011010 will begin other 64-bit instructions,
    0110110 will begin an 80-bit instruction,
    and so on.

    But there's only one container defined so far, the 64-bit container.
    It will contain 18-bit and 36-bit instructions, and others, in 20-bit
    slots.

    The crazy idea for returning from subroutines... means that if the
    full instruction set, including the extra-long instructions that go
    inside the 64-bit container, is used, then the instruction stream can
    switch to containers when the supplementary instructions appear.
    The compiler won't have to do anything special, because branching to instructions inside a container works in every case just like for
    ordinary instructions - you have to use special branching
    instructions, yes, but since one knows what one's destination is,
    that's not a real difficulty.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Stefan Monnier@monnier@iro.umontreal.ca to comp.arch on Thu Sep 24 14:59:31 2026
    From Newsgroup: comp.arch

    John Savard [2026-09-24 04:48:02] wrote:
    Some machines do have less than a gigabyte of memory, and some
    programs are smaller than a gigabyte in size. So if such programs used
    32-bit addressing, they would save memory, and save on time and/or
    power consumption when doing address arithmetic. And there are
    existing architectures with 32-bit address modes as legacy cruft if
    nothing else, with which I would have to compete, and so _if_ the lack
    of it were a problem, _then_ perhaps I should avoid it.

    AFAIK the above does not justify a 32bit architecture. It justifies
    only to make sure the architecture can be used such that all the address
    bits above the 32nd bit are zero, like in the [x32 ABI](https://en.wikipedia.org/wiki/X32_ABI). Most(all?) 64bit
    architectures support that without any special effort, AFAIK.


    === Stefan
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Fri Sep 25 22:29:02 2026
    From Newsgroup: comp.arch

    On Thu, 24 Sep 2026 17:18:03 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:
    quadibloc@invalid.com (John Savard) posted:

    So I am engaging in insane schemes basically for the purpose of
    cheating the Pigeonhole Principle.

    You are also letting perfection get in the way of good enough.

    In any case, I've now come to the conclusion that the proposal on
    which this thread is based is not worth fleshing out,

    Now I think I see how I can squeeze everything I want - which to me
    seems like good enough and not perfection - into something that avoids
    too much weirdness.

    I go back to using all seven base registers for 16-bit displacements,
    instead of going to three to shave off a bit.

    Then, 110 begins a 64-bit or longer instruction,
    111 continues such an instruction, and means do not commence decoding
    on this 32-bit spot.

    So now I've gained - all instructions start out as 32 bits, decoding
    is initially parallel, because every 32 bits say what is to be done
    with it.

    But the loss is that the best I can do with this is have a 64-bit
    container instruction with three 19-bit slots in it. No longer can I
    have 20-bit slots.

    But that's fine, a 19-bit slot will do, as follows:

    0: A short instruction, with 18 remaining bits available.
    10: The first 17 bits of a longer instruction.
    11: An additional 17 bits of a longer instruction.

    Since I'm still using container instructions, instead of block
    structure, I still potentially have the subroutine return issue. I
    could just return from a short instruction stream using a physical
    16-bit aligned address, and if it is in 32 bits that start with 11,
    just backtrack until a 10 is encountered to decode everything.

    But instead I could use the address of the start of the container
    instruction, because now I have two unused bits at the least
    significant end. Those two bits could contain 1, 2, or 3 to indicate
    which embedded instruction is the target.

    A small price is paid. 34 bits is not long enough to handle everything
    I might want to include in "almost" 32 bits. It won't quite fit

    opcode: 7 bits
    C bit to indicate the condition code can be affected: 1 bit
    destination register: 5 bits
    index register: 3 bits
    base register: 3 bits
    displacement: 16 bits

    That totals to 35 bits.

    Memory-to-register operate instructions, though, have a much lower
    importance than getting the basic load-store instructions right, and
    having usable short instructions.

    I no longer have any true 16-bit instructions; my short instructions,
    18 bits long, take up 21 1/3 bits of space each.

    But that's a small price to pay to get fast decoding without getting
    weird and exotic.

    The big price is that now long instructions which include immediate
    values no longer can just include them in the exact same form as data.

    Thus, an instruction with a 64-bit immediate would look like this:

    110: 3 bits
    instruction: 23 bits
    bits 0-2 of the immediate value: 3 bits
    bits 32-34 of the immediate value: 3 bits

    111: 3 bits
    bits 3-31 of the immediate value: 29 bits

    111: 3 bits
    bits 35-63 of the immediate value: 29 bits

    Bit numbering is big-endian from zero.

    The container with three 19-bit slots follows a similar pattern,
    looking like this:

    110: 3 bits
    1: only one bit left to specify a container
    Six-bit prefix to the second short instruction slot: 6 bits
    Three-bit prefix to the third short instruction slot: 3 bits
    First short instruction slot: 19 bits

    111: 3 bits
    Last thirteen bits of the second short instruction slot: 13 bits
    Last sixteen bits of the third short instruction slot: 16 bits

    The idea is to keep as many bits aligned as possible.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Sun Sep 27 15:12:01 2026
    From Newsgroup: comp.arch

    On Fri, 25 Sep 2026 22:29:02 GMT, quadibloc@invalid.com (John Savard)
    wrote:

    A small price is paid. 34 bits is not long enough to handle everything
    I might want to include in "almost" 32 bits. It won't quite fit

    opcode: 7 bits
    C bit to indicate the condition code can be affected: 1 bit
    destination register: 5 bits
    index register: 3 bits
    base register: 3 bits
    displacement: 16 bits

    That totals to 35 bits.

    So I take out the C bit, and it fits in 34 bits. Since the opcode
    doesn't begin with 11, I still have 25% of the opcode space available.

    But

    opcode: 7 bits
    C bit to indicate the condition code can be affected: 1 bit
    destination register: 3 bits
    index register: 3 bits
    base register: 3 bits
    displacement: 16 bits

    still takes up 33 bits, so how do I offer the option of affecting the
    condition codes at least if one of the first eight registers is the destination?

    The solution, of course, is obvious. Since not changing the condition
    codes is already covered for all the registers, I don't need a C bit
    here, it's redundant.

    So, now, after the 10 prefix for "first 17 bits of a long
    instruction", I have to start with 1111 and a length indication for
    the 51-bit and 68-bit and so on instructions within encapsulation.

    It's still livable.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Sun Sep 27 23:59:20 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:

    On Fri, 25 Sep 2026 22:29:02 GMT, quadibloc@invalid.com (John Savard)
    wrote:

    A small price is paid. 34 bits is not long enough to handle everything
    I might want to include in "almost" 32 bits. It won't quite fit

    opcode: 7 bits
    C bit to indicate the condition code can be affected: 1 bit
    destination register: 5 bits
    index register: 3 bits
    base register: 3 bits
    displacement: 16 bits

    That totals to 35 bits.

    So I take out the C bit, and it fits in 34 bits. Since the opcode
    doesn't begin with 11, I still have 25% of the opcode space available.

    But

    opcode: 7 bits
    C bit to indicate the condition code can be affected: 1 bit
    destination register: 3 bits
    index register: 3 bits
    base register: 3 bits
    displacement: 16 bits

    still takes up 33 bits, so how do I offer the option of affecting the condition codes at least if one of the first eight registers is the destination?

    The solution, of course, is obvious. Since not changing the condition
    codes is already covered for all the registers, I don't need a C bit
    here, it's redundant.

    You do not need condition codes.

    So, now, after the 10 prefix for "first 17 bits of a long
    instruction", I have to start with 1111 and a length indication for
    the 51-bit and 68-bit and so on instructions within encapsulation.

    It's still livable.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Mon Sep 28 04:07:07 2026
    From Newsgroup: comp.arch

    On Sun, 27 Sep 2026 23:59:20 GMT, MitchAlsup
    <user5857@newsgrouper.org.invalid> wrote:

    You do not need condition codes.

    I am aware that some modern architectures exist which don't have
    condition codes.

    I am not, however, familiar enough with those architectures to know
    how they manage to work without them. I presume they _don't_ do it by
    means of ancient and obsolete stratagems, like "add and skip if carry"
    where the next instruction is a jump instruction, and combining
    operate instruction and conditional jump operations in the same
    instruction sounds like a recipe for very long instructions.

    I suppose I will have to learn about these architectures; I had felt
    that since the System/360 and the Motorola 68000 and the Power PC were
    good enough for younger me, not just good enough for my grandpappy,
    they ought to be good enough for anyone...

    Of course, when I was brought home from the delivery room, computers
    already had hardware floating-point units and reliable random-access
    memories.

    Mind you, I was almost two years old before a high-level language
    which produced efficient object code was available for the computer
    that made both of those features widely available. And, for those who
    haven't guessed, although the transistor had been invented by that
    time, they were still too expensive for ordinary computers, and so the
    computer I'm talking about, the IBM 704, had to get by with vacuum
    tubes.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Mon Sep 28 04:27:05 2026
    From Newsgroup: comp.arch

    On Mon, 28 Sep 2026 04:07:07 GMT, quadibloc@invalid.com (John Savard)
    wrote:

    On Sun, 27 Sep 2026 23:59:20 GMT, MitchAlsup ><user5857@newsgrouper.org.invalid> wrote:

    You do not need condition codes.

    I am aware that some modern architectures exist which don't have
    condition codes.

    I am not, however, familiar enough with those architectures to know
    how they manage to work without them.

    I have now glanced at two of the main examples, the DEC Alpha and the
    MIPS. What they have are instructions that combine a comparison and a conditional jump based on the result of the comparison.

    All right, then, how do you do multi-precision arithmetic without a
    carry flag? Leave the carry in the previous register, if you use a
    special add instruction?

    The chief evil of condition codes, not being able to put other
    instructions between the one that set them, and the one that tests
    them, was remedied on many RISC architectures by putting a bit in
    operate instructions to turn setting the condition codes on or off. I
    do that.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 28 06:20:43 2026
    From Newsgroup: comp.arch

    quadibloc@invalid.com (John Savard) writes:
    I have now glanced at two of the main examples, the DEC Alpha and the
    MIPS. What they have are instructions that combine a comparison and a >conditional jump based on the result of the comparison.

    MIPS only has that for branch on equality. Alpha only has this for
    comparisons with 0. RISC-V has such instructions indeed.

    What Alpha does have is comparison instructions that produce 0 or 1 in
    a GPR. MIPS has that for inequality; apparently if you want to get a
    0 or 1 for an equality comparison, you need to sythesize it in some
    other way (probably with xor and sltiu, or so). RISC-V has the same
    comparison instructions as MIPS.

    All right, then, how do you do multi-precision arithmetic without a
    carry flag? Leave the carry in the previous register, if you use a
    special add instruction?

    Read

    @InProceedings{ertl25-carry2,
    author = {M. Anton Ertl},
    title = {Multi-precision integer arithmetics},
    booktitle = {Tagungsband des Jahrestreffens 2025 der
    GI-Fachgruppe ``Programmiersprachen und
    Rechenkonzepte''},
    year = {2025},
    series = {INSIGHTS --- Schriftenreihe der Fakult\"at Technik},
    pages = {15--25},
    url = {https://www.complang.tuwien.ac.at/papers/ertl25-carry2.pdf},
    slides-url = {https://www.complang.tuwien.ac.at/papers/ertl25-carry2-slides.pdf},
    url-proceedings = {https://www.dhbw-stuttgart.de/fileadmin/dateien/Forschung/Forschungsschwerpunkte_Technik/DHBW_Stuttgart_INSIGHTS_1_2025_Tagungsband_Jahrestreffen_GI-Fachgruppe_Programmiersprachen_und_Rechenkonzepte_2025.pdf},
    abstract = {Multi-precision integer arithmetics is widely used,
    among other things in public-key cryptography and
    when computing many digits of transcendental
    numbers. The present paper discusses
    multi-precision addition and multiplication:
    architectural support and its use in hand-written
    assembly language, libraries that use such
    assembly-language code, and programming language
    support and how close the code generated by Clang
    and GCC is to the hand-written assembly language.}
    }

    You may also be interested in the unpublished <https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf>.

    The chief evil of condition codes, not being able to put other
    instructions between the one that set them, and the one that tests
    them, was remedied on many RISC architectures by putting a bit in
    operate instructions to turn setting the condition codes on or off. I
    do that.

    Ironically, modern (since around 2010 or so) microarchitectures tend
    to combine compare and branch instructions into one microinstruction
    (now called macroinstruction), and putting other instructions in
    between can be counterproductive.

    The chief evil of condition codes is that on many architectures there
    is only one set of them, architecturally. So you do not have concepts
    in programming languages that correspond to them, and if you have them
    (such as clang), the code they compile them to is not great, because
    you cannot properly register-allocate them.

    Admittedly, programming languages have not contained concepts that
    correspond with condition codes. The main usage is for conditional
    branching, and there the condition codes are just used to transmit
    info between comparison instructions and branch instructions, and one
    set is enough. IIRC Algol 60 does not have a boolean data type (the
    results of comparisons are syntactically separated from arithmetic,
    not by the type system), and I expect that the same is true in early
    Fortran. We only get reified comparison results with the next
    generation of mainstream programming langyages, Pascal and C; and
    there they are rarely used.

    Multi-precision integer arithmetic is still very niche in programming languages. A number of them have big integers, but that's mainly a
    good way to avoid having to deal with integer overflows, not because
    they handle a lot of multi-precision arithmetics (big integers tend to
    be slow for that usage). C23 has added _BitInt types, where gcc and
    clang support 64K and 8M bits, respectively, so there you finally have multi-precision arithmetics as programming language feature.

    And in computer architecture, it is somewhat niche, too.
    Architectures like S/360 do not support it well (the IBM 704 is better
    in that respect), and only ESA390 (IIRC) got addition with carry-in
    and carry-out. Similarly, MIPS, Alpha, and RISC-V make
    multi-precision arithmetics a relatively expensive affair; apparently
    they do not consider that important for market success. The fact that
    Intel ADX is not included in x86-64-v4 despite being available since
    Haswell also indicates that multi-precision arithmetics is not
    considered important by many; is there a CPU that satisfies x86-64-v4
    that does not have ADX?

    PowerPC has 8 sets of condition codes (<,=,>, and IIRC Overflow, not
    the usual set of NCZV), with comparison instructions being able to
    write to any of them, but other instructions can only write to CC0.
    And PowerPC only has one carry flag.

    For multi-precision arithmetics, having more than one carry flag is
    very useful; that's obvious for a computation like a+b+c, and less
    obvious for a computation like a*b. Intel has added the ADX
    extension, which turns the O flag into a second carry flag for more
    efficient multi-precision multiplication and squaring (used in
    cryptography).

    In implementation, condition codes are relatively expensive on OoO
    machines: They have to be renamed separately, as one set on many
    architectures, but three sets on IA-32/AMD64 (ZNP, C, and O); AMD64 implementations tend to have as many rename registers for each of
    these as GPRs, ARM A64 implementations about half as many.

    With carry and overflow as part of the result register, they would
    just cost the actual bits, not all the renamer resources and
    checkpoint resources that more renamer state also costs.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From quadibloc@quadibloc@invalid.com (John Savard) to comp.arch on Mon Sep 28 11:09:51 2026
    From Newsgroup: comp.arch

    On Mon, 28 Sep 2026 06:20:43 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    You may also be interested in the unpublished ><https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf>.

    I am aware that for extremely long numbers, there's a fast
    multiplication technique using a Fast Fourier Transform.

    With carry and overflow as part of the result register, they would
    just cost the actual bits, not all the renamer resources and
    checkpoint resources that more renamer state also costs.

    Given that variable types are matched in width to the memory, this
    would result in register sizes not being matched in width to the
    memory, complicating saves and restores.

    I have tended to include in my architecture optional special forms of
    the operate and conditional jump instructions that make use of an
    alternate set of eight condition codes, to reproduce the capability of
    the Power PC.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 28 17:00:36 2026
    From Newsgroup: comp.arch

    quadibloc@invalid.com (John Savard) writes:
    On Mon, 28 Sep 2026 06:20:43 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    You may also be interested in the unpublished >><https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf>.
    ...
    With carry and overflow as part of the result register, they would
    just cost the actual bits, not all the renamer resources and
    checkpoint resources that more renamer state also costs.

    Given that variable types are matched in width to the memory, this
    would result in register sizes not being matched in width to the
    memory, complicating saves and restores.

    On a regular save or restore, e.g., in the calling convention, the
    flags are typically not saved and not restored. Likewise, ordinary
    store instructions don't store the carry and overflow bits, and
    ordinary load instructions clear the carry and overflow bits. Already
    the IBM 704 dealt in that way with the additional bits of its
    accumulator.

    For context switches, one wants to save and restore the extra bits,
    and this causes a complication, as discussed in the paper. I think
    that the benefits outweigh this complication.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Mon Sep 28 17:13:33 2026
    From Newsgroup: comp.arch

    John Savard <quadibloc@invalid.com> schrieb:
    On Mon, 28 Sep 2026 06:20:43 GMT, anton@mips.complang.tuwien.ac.at
    (Anton Ertl) wrote:

    You may also be interested in the unpublished >><https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf>.

    I am aware that for extremely long numbers, there's a fast
    multiplication technique using a Fast Fourier Transform.

    With carry and overflow as part of the result register, they would
    just cost the actual bits, not all the renamer resources and
    checkpoint resources that more renamer state also costs.

    Given that variable types are matched in width to the memory, this
    would result in register sizes not being matched in width to the
    memory, complicating saves and restores.

    I have tended to include in my architecture optional special forms of
    the operate and conditional jump instructions that make use of an
    alternate set of eight condition codes, to reproduce the capability of
    the Power PC.

    POWER has several condition code registers (eight!), but most of
    them are never used in actual code, as disassembly will show.
    The arithmetic + register instructions (add. and friends)
    instructions only affect CR0.

    If you add this kind of thing, use fewer registers (at most four)
    and put enough bits in them that you can express every condition
    that you might want.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Michael S@already5chosen@yahoo.com to comp.arch on Mon Sep 28 20:23:36 2026
    From Newsgroup: comp.arch

    On Mon, 28 Sep 2026 06:20:43 GMT
    anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

    quadibloc@invalid.com (John Savard) writes:
    I have now glanced at two of the main examples, the DEC Alpha and the
    MIPS. What they have are instructions that combine a comparison and a >conditional jump based on the result of the comparison.

    MIPS only has that for branch on equality.

    It depends on what you consider MIPS.
    Rev.6 provides much more than branch on equality.
    Specifically:
    Compact Branch if Equal
    Compact Branch if Not Equal
    Compact Branch if Greater than or Equal, signed
    Compact Branch if Less Than, signed
    Compact Branch if Greater than or Equal, unsigned
    Compact Branch if Less Than, unsigned
    Compact Branch if Overflow (word)
    Compact Branch if No overflow, word


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 28 18:01:06 2026
    From Newsgroup: comp.arch

    Thomas Koenig <tkoenig@netcologne.de> writes:
    POWER has several condition code registers (eight!), but most of
    them are never used in actual code, as disassembly will show.

    I actually expected that, too, but when I checked it, I found:

    $ objdump -d gforth-fast|grep cr[0-7]
    100040fc: 00 00 a0 2f cmpdi cr7,r0,0
    10004100: 0c 00 fe 41 beq+ cr7,1000410c <_init+0x2c>
    1000699c: 00 00 88 2f cmpwi cr7,r8,0
    10006a30: 00 00 87 2f cmpwi cr7,r7,0
    10008df8: 00 00 1f fc fcmpu cr0,f31,f0
    10008e28: 00 00 1f fc fcmpu cr0,f31,f0
    10008e5c: 00 f8 00 fc fcmpu cr0,f0,f31
    10008e8c: 00 f8 00 fc fcmpu cr0,f0,f31
    10008ebc: 00 f8 00 fc fcmpu cr0,f0,f31
    10008ef0: 00 f8 00 fc fcmpu cr0,f0,f31
    10008f1c: 00 f0 1f fc fcmpu cr0,f31,f30
    10008f4c: 00 f0 1f fc fcmpu cr0,f31,f30
    10008f80: 00 f0 1f fc fcmpu cr0,f31,f30
    10008fac: 00 f0 1f fc fcmpu cr0,f31,f30
    10008fdc: 00 f0 1f fc fcmpu cr0,f31,f30
    1000900c: 00 f0 1f fc fcmpu cr0,f31,f30
    100096dc: 00 f0 1f fc fcmpu cr0,f31,f30
    1000a944: 00 f8 00 fc fcmpu cr0,f0,f31
    1000a980: 00 00 1f fc fcmpu cr0,f31,f0
    100117ac: 00 00 88 2f cmpwi cr7,r8,0
    10011860: 00 00 87 2f cmpwi cr7,r7,0
    100150e8: 00 00 1f fc fcmpu cr0,f31,f0
    10015138: 00 00 1f fc fcmpu cr0,f31,f0
    1001518c: 00 f8 00 fc fcmpu cr0,f0,f31
    100151dc: 00 f8 00 fc fcmpu cr0,f0,f31
    1001522c: 00 f8 00 fc fcmpu cr0,f0,f31
    10015280: 00 f8 00 fc fcmpu cr0,f0,f31
    100152cc: 00 f0 1f fc fcmpu cr0,f31,f30
    1001531c: 00 f0 1f fc fcmpu cr0,f31,f30
    10015370: 00 f0 1f fc fcmpu cr0,f31,f30
    100153bc: 00 f0 1f fc fcmpu cr0,f31,f30
    1001540c: 00 f0 1f fc fcmpu cr0,f31,f30
    1001545c: 00 f0 1f fc fcmpu cr0,f31,f30
    1001616c: 00 f0 1f fc fcmpu cr0,f31,f30
    100182d4: 00 f8 00 fc fcmpu cr0,f0,f31
    10018330: 00 00 1f fc fcmpu cr0,f31,f0
    10021ba0: 00 00 89 2f cmpwi cr7,r9,0
    10021bac: e4 00 9e 40 bne cr7,10021c90 <gforth_go+0x170>
    10021be8: 00 00 89 2f cmpwi cr7,r9,0
    10022c48: 00 00 88 2f cmpwi cr7,r8,0
    10022c54: 1c 01 9e 41 beq cr7,10022d70 <append_ip_update+0x160>
    10022c8c: 54 01 9d 40 ble cr7,10022de0 <append_ip_update+0x1d0>
    100235f4: 00 00 08 2e cmpwi cr4,r8,0
    100235fc: b4 02 92 41 beq cr4,100238b0 <compile_prim_dyn+0x3d0>
    100236a8: c8 00 92 41 beq cr4,10023770 <compile_prim_dyn+0x290>
    10023898: 00 00 08 2e cmpwi cr4,r8,0
    1002389c: 68 fd 92 40 bne cr4,10023604 <compile_prim_dyn+0x124>
    1002393c: 00 00 08 2e cmpwi cr4,r8,0
    10023974: d4 fc 92 40 bne cr4,10023648 <compile_prim_dyn+0x168>
    10023ad0: 00 00 05 2e cmpwi cr4,r5,0
    10023afc: 44 03 91 40 ble cr4,10023e40 <optimize_rewrite.constprop.0+0x490>
    10023e68: dc 08 91 40 ble cr4,10024744 <optimize_rewrite.constprop.0+0xd94>
    10023f7c: d8 07 91 40 ble cr4,10024754 <optimize_rewrite.constprop.0+0xda4>
    1002402c: 00 00 10 2e cmpwi cr4,r16,0
    10024030: 20 03 92 40 bne cr4,10024350 <optimize_rewrite.constprop.0+0x9a0>
    100240e4: 00 00 a8 2f cmpdi cr7,r8,0
    100240e8: 18 02 9e 41 beq cr7,10024300 <optimize_rewrite.constprop.0+0x950>
    100240ec: 0c 00 92 40 bne cr4,100240f8 <optimize_rewrite.constprop.0+0x748>
    1002439c: 50 00 91 40 ble cr4,100243ec <optimize_rewrite.constprop.0+0xa3c>
    1002570c: 01 00 32 2a cmpldi cr4,r18,1
    100257bc: a4 00 92 41 beq cr4,10025860 <check_prims+0x430>
    10025888: 00 00 87 2f cmpwi cr7,r7,0
    100258c0: 90 00 9e 41 beq cr7,10025950 <check_prims+0x520>
    1002594c: 00 00 89 2f cmpwi cr7,r9,0
    10025a28: b4 02 9e 40 bne cr7,10025cdc <check_prims+0x8ac>
    10025a4c: 00 00 88 2f cmpwi cr7,r8,0
    10025a58: 1c 00 9e 41 beq cr7,10025a74 <check_prims+0x644>
    10025ad0: 98 02 9a 41 beq cr6,10025d68 <check_prims+0x938>
    10025b78: 08 01 9e 40 bne cr7,10025c80 <check_prims+0x850>
    10025bc0: e4 00 9e 40 bne cr7,10025ca4 <check_prims+0x874>
    10025c00: 7c ff 9e 41 beq cr7,10025b7c <check_prims+0x74c>
    10025c50: 2c ff 9e 41 beq cr7,10025b7c <check_prims+0x74c>
    10025cd4: 00 00 88 2f cmpwi cr7,r8,0
    10025d7c: 58 fd 9a 40 bne cr6,10025ad4 <check_prims+0x6a4>
    10025d94: 40 fd 9a 40 bne cr6,10025ad4 <check_prims+0x6a4>
    10025dac: 28 fd 9a 40 bne cr6,10025ad4 <check_prims+0x6a4>
    10025eec: 40 58 aa 7f cmpld cr7,r10,r11
    10025ef4: 08 00 9c 41 blt cr7,10025efc <compile_prim1+0xfc>
    10027380: 3a 00 8a 2f cmpwi cr7,r10,58
    100281d0: 40 28 a7 7f cmpld cr7,r7,r5
    100281d8: 38 02 9c 41 blt cr7,10028410 <gforth_printmetrics+0x2e0>
    10028cac: 01 00 8a 2f cmpwi cr7,r10,1
    10028cc8: 01 00 8a 2f cmpwi cr7,r10,1
    1002a6ec: 2f 00 88 2f cmpwi cr7,r8,47
    1002a6f0: 10 02 9e 41 beq cr7,1002a900 <tilde_cstr.part.0+0x260>
    1002b09c: 00 00 a4 2f cmpdi cr7,r4,0
    1002b0cc: 58 00 9e 41 beq cr7,1002b124 <listlfind+0x94>
    1002b154: 00 00 a4 2f cmpdi cr7,r4,0
    1002b190: 20 00 9e 4d beqlr cr7
    1002ba78: 0d 00 a9 2f cmpdi cr7,r9,13
    1002ba7c: 40 f0 ba 7e cmpld cr5,r26,r30
    1002ba84: 70 00 9e 41 beq cr7,1002baf4 <read_line+0x154>
    1002ba90: c0 ff 96 40 bne cr5,1002ba50 <read_line+0xb0>
    1002be2c: 00 08 01 fc fcmpu cr0,f1,f1
    1002be74: 00 60 00 fc fcmpu cr0,f0,f12
    1002ce10: 00 00 25 2e cmpdi cr4,r5,0
    1002ce1c: 10 00 90 40 bge cr4,1002ce2c <fmdiv+0x3c>
    1002ce50: 08 00 90 40 bge cr4,1002ce58 <fmdiv+0x68>
    1002d1e8: 00 60 01 fc fcmpu cr0,f1,f12
    1002d2e8: 00 00 8a 2e cmpwi cr5,r10,0
    1002d2f0: 40 00 a8 2b cmpldi cr7,r8,64
    1002d310: 70 00 9e 41 beq cr7,1002d380 <__floattidf+0xe0>
    1002d3bc: 05 00 9f 42 bcl 20,4*cr7+so,1002d3c0 <__glink_PLTresolve+0x8>

    cr1, cr2, and cr3 are not used in this binary. I wonder why gcc used
    so many of the others.

    If you add this kind of thing, use fewer registers (at most four)
    and put enough bits in them that you can express every condition
    that you might want.

    Power(PC) stores the carry separately, so this one case that would
    benefit from having several instances has only one.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.arch on Mon Sep 28 18:09:08 2026
    From Newsgroup: comp.arch

    Michael S <already5chosen@yahoo.com> writes:
    On Mon, 28 Sep 2026 06:20:43 GMT
    anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

    quadibloc@invalid.com (John Savard) writes:
    I have now glanced at two of the main examples, the DEC Alpha and the
    MIPS. What they have are instructions that combine a comparison and a
    conditional jump based on the result of the comparison.

    MIPS only has that for branch on equality.

    It depends on what you consider MIPS.

    MIPS-I...MIPS-IV.

    Rev.6 provides much more than branch on equality.

    Yes, they made some good changes there, but soon after canceled the architecture. I never came across a machine with that architecture,
    while I came across MIPS-I and MIPS-IV machines.

    - anton
    --
    'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
    Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From MitchAlsup@user5857@newsgrouper.org.invalid to comp.arch on Mon Sep 28 18:19:21 2026
    From Newsgroup: comp.arch


    quadibloc@invalid.com (John Savard) posted:

    On Sun, 27 Sep 2026 23:59:20 GMT, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:

    You do not need condition codes.

    I am aware that some modern architectures exist which don't have
    condition codes.

    I am not, however, familiar enough with those architectures to know
    how they manage to work without them. I presume they _don't_ do it by
    means of ancient and obsolete stratagems, like "add and skip if carry"
    where the next instruction is a jump instruction, and combining
    operate instruction and conditional jump operations in the same
    instruction sounds like a recipe for very long instructions.

    There is a branch on bit instruction
    There is a CMP instruction that sets/clears bits based on the comparison.
    My 66000 integer CMP sets/clears 21 bits, FP sets/clears 31.

    Conditions are simply a vector of bits in a register. You get as many
    as you like.

    I suppose I will have to learn about these architectures; I had felt
    that since the System/360 and the Motorola 68000 and the Power PC were
    good enough for younger me, not just good enough for my grandpappy,
    they ought to be good enough for anyone...

    Of course, when I was brought home from the delivery room, computers
    already had hardware floating-point units and reliable random-access memories.

    Mind you, I was almost two years old before a high-level language
    which produced efficient object code was available for the computer
    that made both of those features widely available. And, for those who
    haven't guessed, although the transistor had been invented by that
    time, they were still too expensive for ordinary computers, and so the computer I'm talking about, the IBM 704, had to get by with vacuum
    tubes.

    John Savard
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Thomas Koenig@tkoenig@netcologne.de to comp.arch on Tue Sep 29 08:06:26 2026
    From Newsgroup: comp.arch

    Anton Ertl <anton@mips.complang.tuwien.ac.at> schrieb:
    Thomas Koenig <tkoenig@netcologne.de> writes:
    POWER has several condition code registers (eight!), but most of
    them are never used in actual code, as disassembly will show.

    I actually expected that, too, but when I checked it, I found:

    [lots of different registers]

    Unusual pattern. I have, for a freshly-compiled f951, from a simple
    grep. Register name, percentage, cumilative percentage, above 100%
    because of register-register moves.

    cr7 20696 49.24-a% 49.24-a%
    cr4 10635 25.30-a% 74.54-a%
    cr6 4879 11.61-a% 86.15-a%
    cr3 2473 5.88-a% 92.03-a%
    cr5 2165 5.15-a% 97.18-a%
    cr2 968 2.30-a% 99.48-a%
    cr0 298 0.71-a% 100.19-a%
    cr1 138 0.33-a% 100.52-a%

    So, this does not look like a lot of register pressure; four
    registers could cover more than 90% of use. This could save a bit
    in the branch instructions, to be used either for wider branches,
    or less opcode space taken up. Adding the carry to the CR
    registers would require a third operand for ADDC etc, soo...
    lots of interesting choices.

    It would be interesting to hack gcc so it only knows about
    four of the condition registers in POWER, and then compare
    generated code, but that would require some backend hackery.
    --
    This USENET posting was made without artificial intelligence,
    artificial impertinence, artificial arrogance, artificial stupidity,
    artificial flavorings or artificial colorants.
    --- Synchronet 3.22a-Linux NewsLink 1.2