I tried to find a way to fit everything I wanted into too little
opcode space. But all the strategems I tried were, rightly, criticized
as being too complex, as being something people would never accept.
So what else is left? In the past, I had tried various ways of shaving
one bit off the length of instructions through various compromises. I
was dissatisfied with them all, and went around in circles for so long
with that.
So I thought I would never go back there. But it seems there's no
other choice.
Thus, I am now going to get close to having everything I wanted - but
not closer than is possible. The scheme is now:
0xx where xx is not 11: 32-bit instructions. With 16-bit
displacements, the base register field will only be two bits long,
allowing three base registers with these instructions.
1: 16-bit instructions. The source and destination registers can only
be registers 0 to 15. This lets them fit.
0110 - 48-bit instructions.
01110 - 64-bit instructions.
Among the 64-bit instructions, there will be an instruction perhaps of
the form 0111011 followed by three 19-bit slots. These slots might
contain short instructions that now can work with all 32 registers for
both operands without restriction, and even set the condition codes.
Or two of those slots could contain a slightly larger than 32-bit
instruction - including memory-reference instructions with all seven
base registers for 32-bit displacements, and memory-to-register
operate instructions.
But because those 64-bit instructions can start at any 16-bit
position, there's a problem. This new "complete" instruction set is
crippled if it is impossible to branch into a sequence of such
instructions except at every third position.
So I need a special branch instruction with an extra two bits. It
probably won't fit in the regular 32-bit instructions, but in the
38-bit instructions, that would be fine.
Oops. One big problem. So the fancy jump instruction has, as its
effective address, the start of the 64-bit instruction containing
19-bit units. Then there's an extra two bits to pick which one. (Only
one bit is "really" needed, since for the first unit, you can just
branch to the 64-bit instruction. Wasting a bit for ease of
comprehension, though, seems appropriate.)
What's the big problem? When do you have *only* an address to
determine where you're branching to, and can't have other fields in
the instruction involved specifying anything else?
When you're making use of the return address given to a subroutine
upon entry!
So here is the insanity I still must engage in.
Return addresses must include special coding. If the return location
is to an instruction that begins within one of these 64-bit capsule instructions... its form must be as follows:
(2 bits): Either 10 or 11, for the second or third 19-bit slot.
Possibly 00 and 01, which are synonymous, for the first one, in the
case of some freak interrupt.
(61 bits): The actual bits specifying the address of the 64-bit
instruction containing those slots.
1: All instructions are multiples of 16 bits in length, and they all
begin on 16-bit boundaries. So odd addresses don't exist; this bit is, therefore, available as the flag to indicate a "special" address that
both points to a 64-bit instruction and to a slot within it.
I think that it's possible to get by with merely 62-bit addressing.
Of course, though, this means that in 32-bit mode, suddenly the
address space is limited to *one* measly gigabyte instead of four! So
in the 32-bit days, this would have been painful indeed.
John Savard--- Synchronet 3.22a-Linux NewsLink 1.2
I think part of the problem is that you continue to lie to yourself.
111yyyyyyyyyyyyy is a 13-bit instruction with a 3-bit marker. >1xxyyyyyyyyyyyyyyyyyyyyyyyyyyyyy is a 29-bit instruction with a 3-bit
marker.
0110 is a 44-bit instruction with a 4-bit marker.
01110 is a 59-bit instruction with a 5-bit marker.
Do not have a CALL instruction that is not on an appropriate boundary,
so, the subsequent instruction is not also on an appropriate boundary.
Why would an architecture, designed in 2026, have a 32-bit address "mode".
Remember:: Nobody else did this to you.
On Wed, 23 Sep 2026 18:14:22 GMT, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
Remember:: Nobody else did this to you.
But am I doing it to myself out of sheer masochism?
No: instead, I'm paying the price for my greed, which is a different situation.
The problem is: I'd like to have an architecture in which the 32-bit instructions do "this", and the 16-bit instructions do "that", with my standards for this and that being based on the IBM System/360.
But the IBM System/360 had 16 general registers (and four
floating-point registers, but in ESA/390 it was easy to expand that to
16 as well). I want to have 32 registers in the integer and floating
register banks because RISC chips that are popular these days have
that.
So I am engaging in insane schemes basically for the purpose of
cheating the Pigeonhole Principle.
It would be simpler and more orthogonal if I were to propose an
architecture which used bit addressing, like the IBM 7030/STRETCH, but
in which, unlike the STRETCH, _instructions_ were arbitrary numbers
of bits in length.
So if memory-reference instructions that can do what I want happen to
take up 37 bits, no problem. They can be 37 bits long.
I presume, though, that such a level of barrel-shifter overuse would
be inefficient. So I'm instead trying to come up with schemes that
play nicely with power-of-two buses and memory.
Thus, I implement instructions that are slightly bigger than 16 bits
and instructions that are slightly bigger than 32 bits... by packing
three 18-bit instruction chunks inside a 64-bit instruction.
The rest is all a consequence of respecting the advice that said that building an ISA around 256-bit instruction blocks is bad, and--- Synchronet 3.22a-Linux NewsLink 1.2
variable-length instructions don't need to be avoided for efficiency.
John Savard
My suggestion is fewer container sizer = easier Decoding and easier
boundary controls.
On Thu, 24 Sep 2026 17:18:03 GMT, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
My suggestion is fewer container sizer = easier Decoding and easier >boundary controls.
So far, I only have one container size.
While a previous proposal did have three 20-bit instructions in 64
bits, now I have to get by with fewer, since I have other instructions
for which to make room.
Seven bits for opcode, one bit to indicate if the condition codes are
set, ten bits for source and destination registers - 18 bits is enough
for a full short instruction. So reserving three 19-bit instruction
slots in 64 bits lets me also have room for 36-bit instructions.
Oh, dear. One bit isn't enough to distinguish between: 18-bit short,
start of 36-bit, and do not decode.
20 bits is possible, I just have to assign fully half of the opcode--- Synchronet 3.22a-Linux NewsLink 1.2
space starting with 011 to these containers. Which is fine; they're important, for making the instruction set extensible.
The container mechanism does allow, since it branches to the container instruction, then indicating an instruction within it, for other
container types. I am envisioning _one_ other container type, with
longer instruction slots, to handle instructions with an explicit
indication of parallelism, and predication.
The idea is to have an ISA that is almost indefinitely extensible - at
the cost of relatively high levels of overhead for the extensions.
John Savard
You stated 5 different instruction container sizes packed into one >fetch-container.
Some machines do have less than a gigabyte of memory, and some
programs are smaller than a gigabyte in size. So if such programs used
32-bit addressing, they would save memory, and save on time and/or
power consumption when doing address arithmetic. And there are
existing architectures with 32-bit address modes as legacy cruft if
nothing else, with which I would have to compete, and so _if_ the lack
of it were a problem, _then_ perhaps I should avoid it.
quadibloc@invalid.com (John Savard) posted:
So I am engaging in insane schemes basically for the purpose of
cheating the Pigeonhole Principle.
You are also letting perfection get in the way of good enough.
A small price is paid. 34 bits is not long enough to handle everything
I might want to include in "almost" 32 bits. It won't quite fit
opcode: 7 bits
C bit to indicate the condition code can be affected: 1 bit
destination register: 5 bits
index register: 3 bits
base register: 3 bits
displacement: 16 bits
That totals to 35 bits.
On Fri, 25 Sep 2026 22:29:02 GMT, quadibloc@invalid.com (John Savard)
wrote:
A small price is paid. 34 bits is not long enough to handle everything
I might want to include in "almost" 32 bits. It won't quite fit
opcode: 7 bits
C bit to indicate the condition code can be affected: 1 bit
destination register: 5 bits
index register: 3 bits
base register: 3 bits
displacement: 16 bits
That totals to 35 bits.
So I take out the C bit, and it fits in 34 bits. Since the opcode
doesn't begin with 11, I still have 25% of the opcode space available.
But
opcode: 7 bits
C bit to indicate the condition code can be affected: 1 bit
destination register: 3 bits
index register: 3 bits
base register: 3 bits
displacement: 16 bits
still takes up 33 bits, so how do I offer the option of affecting the condition codes at least if one of the first eight registers is the destination?
The solution, of course, is obvious. Since not changing the condition
codes is already covered for all the registers, I don't need a C bit
here, it's redundant.
So, now, after the 10 prefix for "first 17 bits of a long--- Synchronet 3.22a-Linux NewsLink 1.2
instruction", I have to start with 1111 and a length indication for
the 51-bit and 68-bit and so on instructions within encapsulation.
It's still livable.
John Savard
You do not need condition codes.
On Sun, 27 Sep 2026 23:59:20 GMT, MitchAlsup ><user5857@newsgrouper.org.invalid> wrote:
You do not need condition codes.
I am aware that some modern architectures exist which don't have
condition codes.
I am not, however, familiar enough with those architectures to know
how they manage to work without them.
I have now glanced at two of the main examples, the DEC Alpha and the
MIPS. What they have are instructions that combine a comparison and a >conditional jump based on the result of the comparison.
All right, then, how do you do multi-precision arithmetic without a
carry flag? Leave the carry in the previous register, if you use a
special add instruction?
The chief evil of condition codes, not being able to put other
instructions between the one that set them, and the one that tests
them, was remedied on many RISC architectures by putting a bit in
operate instructions to turn setting the condition codes on or off. I
do that.
You may also be interested in the unpublished ><https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf>.
With carry and overflow as part of the result register, they would
just cost the actual bits, not all the renamer resources and
checkpoint resources that more renamer state also costs.
On Mon, 28 Sep 2026 06:20:43 GMT, anton@mips.complang.tuwien.ac.at...
(Anton Ertl) wrote:
You may also be interested in the unpublished >><https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf>.
With carry and overflow as part of the result register, they would
just cost the actual bits, not all the renamer resources and
checkpoint resources that more renamer state also costs.
Given that variable types are matched in width to the memory, this
would result in register sizes not being matched in width to the
memory, complicating saves and restores.
On Mon, 28 Sep 2026 06:20:43 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:
You may also be interested in the unpublished >><https://www.complang.tuwien.ac.at/anton/tmp/carry.pdf>.
I am aware that for extremely long numbers, there's a fast
multiplication technique using a Fast Fourier Transform.
With carry and overflow as part of the result register, they would
just cost the actual bits, not all the renamer resources and
checkpoint resources that more renamer state also costs.
Given that variable types are matched in width to the memory, this
would result in register sizes not being matched in width to the
memory, complicating saves and restores.
I have tended to include in my architecture optional special forms of
the operate and conditional jump instructions that make use of an
alternate set of eight condition codes, to reproduce the capability of
the Power PC.
quadibloc@invalid.com (John Savard) writes:
I have now glanced at two of the main examples, the DEC Alpha and the
MIPS. What they have are instructions that combine a comparison and a >conditional jump based on the result of the comparison.
MIPS only has that for branch on equality.
POWER has several condition code registers (eight!), but most of
them are never used in actual code, as disassembly will show.
If you add this kind of thing, use fewer registers (at most four)
and put enough bits in them that you can express every condition
that you might want.
On Mon, 28 Sep 2026 06:20:43 GMT
anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:
quadibloc@invalid.com (John Savard) writes:
I have now glanced at two of the main examples, the DEC Alpha and the
MIPS. What they have are instructions that combine a comparison and a
conditional jump based on the result of the comparison.
MIPS only has that for branch on equality.
It depends on what you consider MIPS.
Rev.6 provides much more than branch on equality.
On Sun, 27 Sep 2026 23:59:20 GMT, MitchAlsup <user5857@newsgrouper.org.invalid> wrote:
You do not need condition codes.
I am aware that some modern architectures exist which don't have
condition codes.
I am not, however, familiar enough with those architectures to know
how they manage to work without them. I presume they _don't_ do it by
means of ancient and obsolete stratagems, like "add and skip if carry"
where the next instruction is a jump instruction, and combining
operate instruction and conditional jump operations in the same
instruction sounds like a recipe for very long instructions.
I suppose I will have to learn about these architectures; I had felt--- Synchronet 3.22a-Linux NewsLink 1.2
that since the System/360 and the Motorola 68000 and the Power PC were
good enough for younger me, not just good enough for my grandpappy,
they ought to be good enough for anyone...
Of course, when I was brought home from the delivery room, computers
already had hardware floating-point units and reliable random-access memories.
Mind you, I was almost two years old before a high-level language
which produced efficient object code was available for the computer
that made both of those features widely available. And, for those who
haven't guessed, although the transistor had been invented by that
time, they were still too expensive for ordinary computers, and so the computer I'm talking about, the IBM 704, had to get by with vacuum
tubes.
John Savard
Thomas Koenig <tkoenig@netcologne.de> writes:
POWER has several condition code registers (eight!), but most of
them are never used in actual code, as disassembly will show.
I actually expected that, too, but when I checked it, I found:
| Sysop: | Amessyroom |
|---|---|
| Location: | Fayetteville, NC |
| Users: | 74 |
| Nodes: | 6 (0 / 6) |
| Uptime: | 02:20:29 |
| Calls: | 1,194 |
| Files: | 1,353 |
| D/L today: |
2 files (1,590K bytes) |
| Messages: | 291,157 |