• Forth on ARM64

    From antispam@antispam@fricas.org (Waldek Hebisch) to comp.lang.forth on Thu Sep 3 20:51:14 2026
    From Newsgroup: comp.lang.forth

    Under Linux on ARM64 trying to set machine stack pointer to
    value which is not divisible by 16 leads to error. AFAICS this
    means that in default setting machine stack pointer is not
    usable as as Forth user stack pointer or return stack pointer.
    I wonder what Forth implementation do? Do they use different
    registers as user and return stack pointer? Maybe they use
    machine stack pointer for control and locals? Or maybe some
    system magic removes the restriction?
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From peter@peter.noreply@tin.it to comp.lang.forth on Fri Sep 4 07:33:11 2026
    From Newsgroup: comp.lang.forth

    On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
    antispam@fricas.org (Waldek Hebisch) wrote:
    Under Linux on ARM64 trying to set machine stack pointer to
    value which is not divisible by 16 leads to error. AFAICS this
    means that in default setting machine stack pointer is not
    usable as as Forth user stack pointer or return stack pointer.
    I wonder what Forth implementation do? Do they use different
    registers as user and return stack pointer? Maybe they use
    machine stack pointer for control and locals? Or maybe some
    system magic removes the restriction?

    Here is the register assignments for the token VM I wrote for
    ARM64 for lxf 64
    /*
    VM8 assembler based aarch64 vm for lxf64 Forth
    Copyright 2020 Peter FElth
    Register usage
    X19 ip vm instruction pointer
    x20 TOP top of stack cached in x20
    x21 sp vm stack pointer
    x22 rp vm return stack pointer
    x23 fp vm floating point stack pointer
    x24 lp vm local stack pointer
    x25 idx loop index of innermost loop
    x26 limit loop limit of innermost loop
    x27 address of jump table
    d8 FTOP top of float stack cached in d8
    sequence to nest to next opcode is RELOAD
    ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
    ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
    br x2 jump to next machine code
    */
    There are just 2 calls in the hole VM, in these cases 2 registers are
    pushed to maintain 16 byte alignment.
    You can avoid the 16 byte alignment by using a register other then sp
    for the processor stack. My tests showed this code to be about 30% slower
    in execution speed.
    BR
    Peter
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From albert@albert@spenarnc.xs4all.nl to comp.lang.forth on Fri Sep 4 10:21:49 2026
    From Newsgroup: comp.lang.forth

    In article <20260904073311.00003526@tin.it>,
    peter <peter.noreply@tin.it> wrote:
    On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
    antispam@fricas.org (Waldek Hebisch) wrote:

    Under Linux on ARM64 trying to set machine stack pointer to
    value which is not divisible by 16 leads to error. AFAICS this
    means that in default setting machine stack pointer is not
    usable as as Forth user stack pointer or return stack pointer.
    I wonder what Forth implementation do? Do they use different
    registers as user and return stack pointer? Maybe they use
    machine stack pointer for control and locals? Or maybe some
    system magic removes the restriction?


    Here is the register assignments for the token VM I wrote for
    ARM64 for lxf 64

    /*
    VM8 assembler based aarch64 vm for lxf64 Forth
    Copyright 2020 Peter FElth

    Register usage
    X19 ip vm instruction pointer
    x20 TOP top of stack cached in x20
    x21 sp vm stack pointer
    x22 rp vm return stack pointer
    x23 fp vm floating point stack pointer
    x24 lp vm local stack pointer
    x25 idx loop index of innermost loop
    x26 limit loop limit of innermost loop
    x27 address of jump table
    d8 FTOP top of float stack cached in d8

    sequence to nest to next opcode is RELOAD

    ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
    ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
    br x2 jump to next machine code

    */

    There are just 2 calls in the hole VM, in these cases 2 registers are
    pushed to maintain 16 byte alignment.

    You can avoid the 16 byte alignment by using a register other then sp
    for the processor stack. My tests showed this code to be about 30% slower
    in execution speed.

    The problem you have is unrelated to linux, but more a c-compatibility.

    With my assembler Forth I have no problem in linux.
    All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc. The stack pointer SP plays no role in the whole Forth, it is arbitrarily
    mapped to R13. I could map SPO to x21 equally well.

    [In thumb there seems to be a an SP but I do not do thumb.]

    There is a similar problem in DLL calls for Windows system.
    I have a brilliant solution, but the margin is to small to
    contain it.

    You could consult :
    https://github.com/albertvanderhorst/ciforth/wiki


    BR
    Peter

    Groetjes Albert


    --
    The Chinese government is satisfied with its military superiority over USA.
    The next 5 year plan has as primary goal to advance life expectancy
    over 80 years, like Western Europe.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From antispam@antispam@fricas.org (Waldek Hebisch) to comp.lang.forth on Fri Sep 4 21:19:37 2026
    From Newsgroup: comp.lang.forth

    albert@spenarnc.xs4all.nl wrote:
    In article <20260904073311.00003526@tin.it>,
    peter <peter.noreply@tin.it> wrote:
    On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
    antispam@fricas.org (Waldek Hebisch) wrote:

    Under Linux on ARM64 trying to set machine stack pointer to
    value which is not divisible by 16 leads to error. AFAICS this
    means that in default setting machine stack pointer is not
    usable as as Forth user stack pointer or return stack pointer.
    I wonder what Forth implementation do? Do they use different
    registers as user and return stack pointer? Maybe they use
    machine stack pointer for control and locals? Or maybe some
    system magic removes the restriction?


    Here is the register assignments for the token VM I wrote for
    ARM64 for lxf 64

    /*
    VM8 assembler based aarch64 vm for lxf64 Forth
    Copyright 2020 Peter F|nlth

    Register usage
    X19 ip vm instruction pointer
    x20 TOP top of stack cached in x20
    x21 sp vm stack pointer
    x22 rp vm return stack pointer
    x23 fp vm floating point stack pointer
    x24 lp vm local stack pointer
    x25 idx loop index of innermost loop
    x26 limit loop limit of innermost loop
    x27 address of jump table
    d8 FTOP top of float stack cached in d8

    sequence to nest to next opcode is RELOAD

    ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
    ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
    br x2 jump to next machine code

    */

    There are just 2 calls in the hole VM, in these cases 2 registers are >>pushed to maintain 16 byte alignment.

    You can avoid the 16 byte alignment by using a register other then sp
    for the processor stack. My tests showed this code to be about 30% slower >>in execution speed.

    The problem you have is unrelated to linux, but more a c-compatibility.

    With my assembler Forth I have no problem in linux.
    All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc. The stack pointer SP plays no role in the whole Forth, it is arbitrarily mapped to R13. I could map SPO to x21 equally well.

    I wrote "machine stack pointer" (or if you prefer register number 31)
    because it is special on ARM64. Of course, Forth can use a different
    register, but not using machine stack pointer is a waste. C
    compatibility makes this waste more painful, because there are
    only 11 registers not touched by C. For traditional Forth
    implementations this is enough, but I am looking at generating
    machine code.
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From peter@peter.noreply@tin.it to comp.lang.forth on Fri Sep 4 23:51:29 2026
    From Newsgroup: comp.lang.forth

    On Fri, 4 Sep 2026 21:19:37 -0000 (UTC)
    antispam@fricas.org (Waldek Hebisch) wrote:
    albert@spenarnc.xs4all.nl wrote:
    In article <20260904073311.00003526@tin.it>,
    peter <peter.noreply@tin.it> wrote:
    On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
    antispam@fricas.org (Waldek Hebisch) wrote:

    Under Linux on ARM64 trying to set machine stack pointer to
    value which is not divisible by 16 leads to error. AFAICS this
    means that in default setting machine stack pointer is not
    usable as as Forth user stack pointer or return stack pointer.
    I wonder what Forth implementation do? Do they use different
    registers as user and return stack pointer? Maybe they use
    machine stack pointer for control and locals? Or maybe some
    system magic removes the restriction?


    Here is the register assignments for the token VM I wrote for
    ARM64 for lxf 64

    /*
    VM8 assembler based aarch64 vm for lxf64 Forth
    Copyright 2020 Peter FElth

    Register usage
    X19 ip vm instruction pointer
    x20 TOP top of stack cached in x20
    x21 sp vm stack pointer
    x22 rp vm return stack pointer
    x23 fp vm floating point stack pointer
    x24 lp vm local stack pointer
    x25 idx loop index of innermost loop
    x26 limit loop limit of innermost loop
    x27 address of jump table
    d8 FTOP top of float stack cached in d8

    sequence to nest to next opcode is RELOAD

    ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
    ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
    br x2 jump to next machine code

    */

    There are just 2 calls in the hole VM, in these cases 2 registers are >>pushed to maintain 16 byte alignment.

    You can avoid the 16 byte alignment by using a register other then sp
    for the processor stack. My tests showed this code to be about 30% slower >>in execution speed.

    The problem you have is unrelated to linux, but more a c-compatibility.

    With my assembler Forth I have no problem in linux.
    All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc.
    The stack pointer SP plays no role in the whole Forth, it is arbitrarily mapped to R13. I could map SPO to x21 equally well.

    I wrote "machine stack pointer" (or if you prefer register number 31)
    because it is special on ARM64. Of course, Forth can use a different register, but not using machine stack pointer is a waste. C
    compatibility makes this waste more painful, because there are
    only 11 registers not touched by C. For traditional Forth
    implementations this is enough, but I am looking at generating
    machine code.

    I will also in the future make a code generator for ARM64.
    I have one for X64 now!
    my idea is to use
    stp x29, x30, [sp, -16]!
    and
    ldp x29, x30, [sp], 16
    at the start and end of words that has call in them
    r r> r@ can be done with
    stp xzr, x20, [sp, -16]!
    and
    ldp xzr, x20, [sp], 16
    This keeps everything aligned at 16 bytes
    I will use all registers for my code generator. To handle C library compatibility I will have a gate that all C calls pass that saves
    needed registers.
    This is how I handle it on X64 now. It works great.
    X64 also has a requirement of 16 byte alignment of the return stack.
    This is not enforced by the CPU, but it can bite you for specific
    opcodes that require 16 byte alignment. It is worse as Windows and Linux
    have different opinions on when the stack should be aligned, before or
    after the call?
    I handle that now by not aligning in code I generate and instead switch to
    a properly aligned stack at the call gate.
    BR
    Peter
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From antispam@antispam@fricas.org (Waldek Hebisch) to comp.lang.forth on Fri Sep 4 22:25:20 2026
    From Newsgroup: comp.lang.forth

    peter <peter.noreply@tin.it> wrote:
    On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
    antispam@fricas.org (Waldek Hebisch) wrote:

    Under Linux on ARM64 trying to set machine stack pointer to
    value which is not divisible by 16 leads to error. AFAICS this
    means that in default setting machine stack pointer is not
    usable as as Forth user stack pointer or return stack pointer.
    I wonder what Forth implementation do? Do they use different
    registers as user and return stack pointer? Maybe they use
    machine stack pointer for control and locals? Or maybe some
    system magic removes the restriction?


    Here is the register assignments for the token VM I wrote for
    ARM64 for lxf 64

    /*
    VM8 assembler based aarch64 vm for lxf64 Forth
    Copyright 2020 Peter F|nlth

    Register usage
    X19 ip vm instruction pointer
    x20 TOP top of stack cached in x20
    x21 sp vm stack pointer
    x22 rp vm return stack pointer
    x23 fp vm floating point stack pointer
    x24 lp vm local stack pointer
    x25 idx loop index of innermost loop
    x26 limit loop limit of innermost loop
    x27 address of jump table
    d8 FTOP top of float stack cached in d8

    sequence to nest to next opcode is RELOAD

    ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
    ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
    br x2 jump to next machine code

    */

    There are just 2 calls in the hole VM, in these cases 2 registers are
    pushed to maintain 16 byte alignment.

    You can avoid the 16 byte alignment by using a register other then sp
    for the processor stack. My tests showed this code to be about 30% slower
    in execution speed.

    What do you compare? Token threaded code to traditional threaded
    code using memory addresses?

    On ARM64 calls have range +-128MB relative. So if the program code
    does not exceed 128MB, then subroutine threaded code is smaller
    than traditional threaded code and probably quite a bit faster.
    Drawback is that non-leaf words need to push and pop return
    address from the machine stack, using 16 bytes of stack space.
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From peter@peter.noreply@tin.it to comp.lang.forth on Sat Sep 5 09:54:19 2026
    From Newsgroup: comp.lang.forth

    On Fri, 4 Sep 2026 22:25:20 -0000 (UTC)
    antispam@fricas.org (Waldek Hebisch) wrote:
    peter <peter.noreply@tin.it> wrote:
    On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
    antispam@fricas.org (Waldek Hebisch) wrote:

    Under Linux on ARM64 trying to set machine stack pointer to
    value which is not divisible by 16 leads to error. AFAICS this
    means that in default setting machine stack pointer is not
    usable as as Forth user stack pointer or return stack pointer.
    I wonder what Forth implementation do? Do they use different
    registers as user and return stack pointer? Maybe they use
    machine stack pointer for control and locals? Or maybe some
    system magic removes the restriction?


    Here is the register assignments for the token VM I wrote for
    ARM64 for lxf 64

    /*
    VM8 assembler based aarch64 vm for lxf64 Forth
    Copyright 2020 Peter FElth

    Register usage
    X19 ip vm instruction pointer
    x20 TOP top of stack cached in x20
    x21 sp vm stack pointer
    x22 rp vm return stack pointer
    x23 fp vm floating point stack pointer
    x24 lp vm local stack pointer
    x25 idx loop index of innermost loop
    x26 limit loop limit of innermost loop
    x27 address of jump table
    d8 FTOP top of float stack cached in d8

    sequence to nest to next opcode is RELOAD

    ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
    ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
    br x2 jump to next machine code

    */

    There are just 2 calls in the hole VM, in these cases 2 registers are pushed to maintain 16 byte alignment.

    You can avoid the 16 byte alignment by using a register other then sp
    for the processor stack. My tests showed this code to be about 30% slower in execution speed.

    What do you compare? Token threaded code to traditional threaded
    code using memory addresses?

    On ARM64 calls have range +-128MB relative. So if the program code
    does not exceed 128MB, then subroutine threaded code is smaller
    than traditional threaded code and probably quite a bit faster.
    Drawback is that non-leaf words need to push and pop return
    address from the machine stack, using 16 bytes of stack space.

    I have rerun some tests on my old RPi4. Now I do not get any difference
    in speed in using sp as stack pointer or another register. I might remember wrong as it was several years ago i did test this last time.
    I test a fib defined as code so threading type should not influence.
    code fib2
    START:
    stp x29, x30, [sp, -16]!
    cmp x20, 2
    b.ge L1
    mov x20, 1
    b END
    L1:
    sub x10, x20, 1
    str x20, [x21, -8]!
    mov x20, x10
    bl START
    ldr x9, [x21]
    sub x9, x9, 2
    str x20, [x21]
    mov x20, x9
    bl START
    ldr x8, [x21], 8
    add x20, x20, x8
    END:
    ldp x29, x30, [sp], 16
    ret
    end-code
    note that this is a wrong fib as the fib in onebench.fs used by gforth have the same problem. I wanted the same number of iterations.
    My system works in 3 steps.
    1. Compile forth source for token threaded VM
    2. Compile the token code to native assembler code
    3. Assemble with a built in assembler to native code
    For ARM64 only the first step is implemented
    Peter
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From albert@albert@spenarnc.xs4all.nl to comp.lang.forth on Sat Sep 5 13:11:54 2026
    From Newsgroup: comp.lang.forth

    In article <117fcl7$26t23$1@paganini.bofh.team>,
    Waldek Hebisch <antispam@fricas.org> wrote:
    albert@spenarnc.xs4all.nl wrote:
    I wrote "machine stack pointer" (or if you prefer register number 31)
    because it is special on ARM64. Of course, Forth can use a different >register, but not using machine stack pointer is a waste. C
    compatibility makes this waste more painful, because there are
    only 11 registers not touched by C. For traditional Forth
    implementations this is enough, but I am looking at generating
    machine code.

    I do not buy that I can use only 11 registers not touched by C.
    I can easily save the 3 registers needed for Forth, if I want
    to call a C-routine. So I freely use all of them.

    Waldek Hebisch
    --
    The Chinese government is satisfied with its military superiority over USA.
    The next 5 year plan has as primary goal to advance life expectancy
    over 80 years, like Western Europe.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Paul Rubin@no.email@nospam.invalid to comp.lang.forth on Sat Sep 5 15:49:17 2026
    From Newsgroup: comp.lang.forth

    antispam@fricas.org (Waldek Hebisch) writes:
    there are only 11 registers not touched by C.

    I missed some parts of this thread but why would a C compiler leave 11 registers untouched? Usually a compiler with register allocation will
    use all the registers, unless you ask it to reserve some for other
    purposes.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From antispam@antispam@fricas.org (Waldek Hebisch) to comp.lang.forth on Sat Sep 5 23:39:29 2026
    From Newsgroup: comp.lang.forth

    Paul Rubin <no.email@nospam.invalid> wrote:
    antispam@fricas.org (Waldek Hebisch) writes:
    there are only 11 registers not touched by C.

    I missed some parts of this thread but why would a C compiler leave 11 registers untouched? Usually a compiler with register allocation will
    use all the registers, unless you ask it to reserve some for other
    purposes.

    What I wrote was a shortcut. Full version is: when you call a routine
    in C, C compiler will ensure that 11 registers are preserved. That
    is calling convention. And yes, unless register is reserved for
    some global use C compiler may use it. Which means that any
    register that is not designated as preserved may be clobbered by
    C code. There is possible exception for global registers like pointer
    to thread data, but changing value of such register is likely to
    cause trouble for C code, so it is unwise to use them for different
    purpose. And on ARM64 there seem to be no such register among general registers.
    --
    Waldek Hebisch
    --- Synchronet 3.22a-Linux NewsLink 1.2