• F>R, FR@ and FR> Float to return stack words

    From peter@peter.noreply@tin.it to comp.lang.forth on Fri Aug 28 00:04:28 2026
    From Newsgroup: comp.lang.forth


    I recently wanted to adapt compelx-kahan.fs, the complex implementation
    by Julian V. Noble and David N. Williams, to the capabilities of
    ntf64/lxf64.

    One thing I wanted to get away from was the Baden-style float locals,
    fa fb .. and to-fa to-fb .. and the use of temporary buffers.

    I replaced them with F>R, FR@ and FR>. Moving floats to the return stack.
    This works well as both stacks are 64 bits. The effect was very good.

    Taking Z* as an example:

    Original Z* using 2 temporary buffers

    : z* (f: x y u v -- x*u-y*v x*v+y*u)
    (
    Uses the algorithm followed by jvn:
    [x+iy]*[u+iv] = [[x+y]*u - y*[u+v]] + i[[x+y]*u + x*[v-u]]
    Requires 3 multiplications and 5 additions.
    )
    zdup f+ (f: x y u v u+v)
    [ fnoname ] literal f! (f: x y u v)
    fover f- (f: x y u v-u)
    [ fnoname float+ ] literal f! (f: x y u)
    frot fdup (f: y u x x)
    [ fnoname float+ ] literal f@ (f: y u x x v-u)
    f*
    [ fnoname float+ ] literal f! (f: y u x)
    frot fdup (f: u x y y)
    [ fnoname ] literal f@ (f: u x y y u+v)
    f*
    [ fnoname ] literal f! (f: u x y)
    f+ f* fdup (f: u*[x+y] u*[x+y])
    [ fnoname ] literal f@ f- (f: u*[x+y] x*u-y*v)
    fswap
    [ fnoname float+ ] literal f@ (f: x*u-y*v u*[x+y] x*[v-u])
    f+ ; (f: x*u-y*v x*v+y*u)

    With more or less direct exchange to using the return stack
    I got

    : z* ( f: x y u v -- x*u-y*v x*v+y*u)
    zdup f+ f>r ( f: x y u v r: u+v)
    fover f- f>r ( f: x y u r: u+v v-u)
    frot fdup fr> ( f: y u x x v-u r: u+v)
    f* ( f: y u x x*'v-u' r: u+v)
    fr> fswap f>r f>r ( f: y u x r: x*'v-u' u+v)
    frot fdup fr> ( f: u x y y u+v r: x*'v-u')
    f* f>r ( f: u x y r: x*'v-u' y*'u+v')
    f+ f* fdup fr> f- fswap fr> ( f: x*u-y*v u*[x+y] x*[v-u])
    f+ ; ( f: x*u-y*v x*v+y*u)

    The original give the following assembler code

    seea Z*
    0x4291A0 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
    0x4291A6 C5FB110DD2965D00 vmovsd qword ptr [rel 0xA02880], xmm1
    0x4291AE C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
    0x4291B4 C5FB1105CC965D00 vmovsd qword ptr [rel 0xA02888], xmm0
    0x4291BC C5FB1005C4965D00 vmovsd xmm0, qword ptr [rel 0xA02888]
    0x4291C4 C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10]
    0x4291CA C5FB1105B6965D00 vmovsd qword ptr [rel 0xA02888], xmm0
    0x4291D2 C5FB1005A6965D00 vmovsd xmm0, qword ptr [rel 0xA02880]
    0x4291DA C4C17B594508 vmulsd xmm0, xmm0, qword ptr [r13+0x8]
    0x4291E0 C5FB110598965D00 vmovsd qword ptr [rel 0xA02880], xmm0
    0x4291E8 C4C17B104508 vmovsd xmm0, qword ptr [r13+0x8]
    0x4291EE C4C17B584510 vaddsd xmm0, xmm0, qword ptr [r13+0x10]
    0x4291F4 C4C17B594500 vmulsd xmm0, xmm0, qword ptr [r13]
    0x4291FA C5FB100D7E965D00 vmovsd xmm1, qword ptr [rel 0xA02880]
    0x429202 C5FB5CD1 vsubsd xmm2, xmm0, xmm1
    0x429206 C5FB100D7A965D00 vmovsd xmm1, qword ptr [rel 0xA02888]
    0x42920E C5F358C8 vaddsd xmm1, xmm1, xmm0
    0x429212 C5F310C1 vmovsd xmm0, xmm1, xmm1
    0x429216 C4C17B115510 vmovsd qword ptr [r13+0x10], xmm2
    0x42921C 4D8D6D10 lea r13, [r13+0x10]
    0x429220 C3 ret
    129 bytes, 21 instructions
    ok

    There are 3 multiplications, 5 add/sub and 8 loads/stores to the buffer

    The revised version gives

    seea Z*
    0x428A50 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
    0x428A56 C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
    0x428A5C C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10]
    0x428A62 C4C173594D08 vmulsd xmm1, xmm1, qword ptr [r13+0x8]
    0x428A68 C4C17B105508 vmovsd xmm2, qword ptr [r13+0x8]
    0x428A6E C4C16B585510 vaddsd xmm2, xmm2, qword ptr [r13+0x10]
    0x428A74 C4C16B595500 vmulsd xmm2, xmm2, qword ptr [r13]
    0x428A7A C5EB5CD9 vsubsd xmm3, xmm2, xmm1
    0x428A7E C5FB58C2 vaddsd xmm0, xmm0, xmm2
    0x428A82 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
    0x428A88 4D8D6D10 lea r13, [r13+0x10]
    0x428A8C C3 ret
    61 bytes, 12 instructions
    ok

    Half the size and all memory loads/stores gone.

    With the returns stack use I can also do things like

    : +ulp f>r r> 1+ >r fr> ; ok

    0e +ulp e. 5e-324 ok
    0e +ulp fh. 0x0.0000000000001p-1022 ok

    seea +ulp
    0x42A260 C4E1F97EC0 vmovq rax, xmm0
    0x42A265 48FFC0 inc rax
    0x42A268 C4E1F96EC0 vmovq xmm0, rax
    0x42A26D C3 ret
    14 bytes, 4 instructions
    ok

    An alternative I considered was to introduce float locals. It would probably have been more complicated. The result, in the best case, would be equal in
    the generated code. The problem with locals is the scope, that is to the end of the word. Of course the source would be more readable!

    I might still implement float locals.

    What do you have in your system?

    BR
    Peter

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From dxf@dxforth@gmail.com to comp.lang.forth on Fri Aug 28 14:44:37 2026
    From Newsgroup: comp.lang.forth

    On 28/08/2026 8:04 am, peter wrote:
    ...
    What do you have in your system?

    Because return-stacking is more complicated, I use fvariable where I can.
    Where I can't, I have the following def's to fall back on. I've not
    played enough with Baden-type flocals to get a sense of what they're like
    in practice. The varying HERE (PAD, HOLD etc) might be an issue.

    \ x86
    code F>R ( r -- )
    1 floats # bp sub bp push ' f! ) jmp
    end-code

    code FR> ( -- r )
    bp push c: f@ ;c 1 floats # bp add next
    end-code


    \ 8080
    code F>R ( r -- )
    rpp lhld -1 floats d lxi d dad rpp shld
    h push ' f! jmp
    end-code

    code FR> ( -- r )
    rpp lhld h push c: f@ ;c rpp lhld
    1 floats d lxi d dad rpp shld next
    end-code

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.lang.forth on Fri Aug 28 06:50:07 2026
    From Newsgroup: comp.lang.forth

    peter <peter.noreply@tin.it> writes:
    : z* ( f: x y u v -- x*u-y*v x*v+y*u)
    zdup f+ f>r ( f: x y u v r: u+v)
    fover f- f>r ( f: x y u r: u+v v-u)
    frot fdup fr> ( f: y u x x v-u r: u+v)
    f* ( f: y u x x*'v-u' r: u+v)
    fswap f>r f>r ( f: y u x r: x*'v-u' u+v)
    frot fdup fr> ( f: u x y y u+v r: x*'v-u')
    f* f>r ( f: u x y r: x*'v-u' y*'u+v')
    f+ f* fdup fr> f- fswap fr> ( f: x*u-y*v u*[x+y] x*[v-u])
    f+ ; ( f: x*u-y*v x*v+y*u)
    [...]
    seea Z*
    0x428A50 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
    0x428A56 C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
    0x428A5C C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10] >0x428A62 C4C173594D08 vmulsd xmm1, xmm1, qword ptr [r13+0x8] >0x428A68 C4C17B105508 vmovsd xmm2, qword ptr [r13+0x8]
    0x428A6E C4C16B585510 vaddsd xmm2, xmm2, qword ptr [r13+0x10] >0x428A74 C4C16B595500 vmulsd xmm2, xmm2, qword ptr [r13]
    0x428A7A C5EB5CD9 vsubsd xmm3, xmm2, xmm1
    0x428A7E C5FB58C2 vaddsd xmm0, xmm0, xmm2
    0x428A82 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
    0x428A88 4D8D6D10 lea r13, [r13+0x10]
    0x428A8C C3 ret
    61 bytes, 12 instructions

    Nice.

    An alternative I considered was to introduce float locals. It would probably >have been more complicated. The result, in the best case, would be equal in >the generated code. The problem with locals is the scope, that is to the end >of the word. Of course the source would be more readable!

    I might still implement float locals.

    What do you have in your system?

    Gforth has had FP locals since 1994. We recently added F>R, FR> and
    FR@. The reason why we did not have them was that the return-stack
    pointer is not necessarily aligned for FP loads and stores, so on many
    hardware platforms of the 1990s F>R and FR> would have been expensive operations.

    With 64-bit cells becoming mainstream, 64-bit floats (what we use in
    Gforth), and the death of most architectures that do not support
    unaligned accesses, the balance has changed in the meantime.

    - anton
    --
    M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
    comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
    New standard: https://forth-standard.org/
    EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From marcel hendrix@mhx@iae.nl to comp.lang.forth on Fri Aug 28 19:35:33 2026
    From Newsgroup: comp.lang.forth

    On 8/28/2026 12:04 AM, peter wrote:

    I recently wanted to adapt compelx-kahan.fs, the complex implementation
    by Julian V. Noble and David N. Williams, to the capabilities of
    ntf64/lxf64.

    One thing I wanted to get away from was the Baden-style float locals,
    fa fb .. and to-fa to-fb .. and the use of temporary buffers.

    I replaced them with F>R, FR@ and FR>. Moving floats to the return stack. This works well as both stacks are 64 bits. The effect was very good.

    ... > What do you have in your system?

    BR
    Peter

    Using the FPU:
    FORTH> see z*
    Flags: ANSI
    $0135F340 : Z*
    $0135F34A fpop,
    $0135F354 sub rsi, #16 b#
    $0135F358 fstp [rsi] tbyte
    $0135F35A fpop,
    $0135F364 sub rsi, #16 b#
    $0135F368 fstp [rsi] tbyte
    $0135F36A fpop,
    $0135F374 sub rsi, #16 b#
    $0135F378 fstp [rsi] tbyte
    $0135F37A fpop,
    $0135F384 sub rsi, #16 b#
    $0135F388 fstp [rsi] tbyte
    $0135F38A fld [rsi] tbyte
    $0135F38C fld [rsi #32 +] tbyte
    $0135F38F fmulp ST(1), ST
    $0135F391 fld [rsi #16 +] tbyte
    $0135F394 fld [rsi #48 +] tbyte
    $0135F397 fmulp ST(1), ST
    $0135F399 fsubp ST(1), ST
    $0135F39B fld [rsi] tbyte
    $0135F39D fld [rsi #48 +] tbyte
    $0135F3A0 fmulp ST(1), ST
    $0135F3A2 fld [rsi #16 +] tbyte
    $0135F3A5 fld [rsi #32 +] tbyte
    $0135F3A8 fmulp ST(1), ST
    $0135F3AA faddp ST(1), ST
    $0135F3AC lea r13, [r13 #-32 +] qword
    $0135F3B0 fxch ST(2)
    $0135F3B2 fstp [r13 #16 +] tbyte
    $0135F3B6 fstp [r13 0 +] tbyte
    $0135F3BA add rsi, #64 b#
    $0135F3BE ;
    116 bytes.

    -marcel

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From peter@peter.noreply@tin.it to comp.lang.forth on Sun Aug 30 10:06:37 2026
    From Newsgroup: comp.lang.forth

    On Fri, 28 Aug 2026 00:04:28 +0200
    peter <peter.noreply@tin.it> wrote:


    I recently wanted to adapt compelx-kahan.fs, the complex implementation
    by Julian V. Noble and David N. Williams, to the capabilities of ntf64/lxf64.

    One thing I wanted to get away from was the Baden-style float locals,
    fa fb .. and to-fa to-fb .. and the use of temporary buffers.

    I replaced them with F>R, FR@ and FR>. Moving floats to the return stack. This works well as both stacks are 64 bits. The effect was very good.



    : z* ( f: x y u v -- x*u-y*v x*v+y*u)
    zdup f+ f>r ( f: x y u v r: u+v)
    fover f- f>r ( f: x y u r: u+v v-u)
    frot fdup fr> ( f: y u x x v-u r: u+v)
    f* ( f: y u x x*'v-u' r: u+v)
    fr> fswap f>r f>r ( f: y u x r: x*'v-u' u+v)
    frot fdup fr> ( f: u x y y u+v r: x*'v-u')
    f* f>r ( f: u x y r: x*'v-u' y*'u+v')
    f+ f* fdup fr> f- fswap fr> ( f: x*u-y*v u*[x+y] x*[v-u])
    f+ ; ( f: x*u-y*v x*v+y*u)



    There are 3 multiplications, 5 add/sub and 8 loads/stores to the buffer

    The revised version gives

    seea Z*
    0x428A50 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
    0x428A56 C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
    0x428A5C C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10] 0x428A62 C4C173594D08 vmulsd xmm1, xmm1, qword ptr [r13+0x8] 0x428A68 C4C17B105508 vmovsd xmm2, qword ptr [r13+0x8]
    0x428A6E C4C16B585510 vaddsd xmm2, xmm2, qword ptr [r13+0x10] 0x428A74 C4C16B595500 vmulsd xmm2, xmm2, qword ptr [r13]
    0x428A7A C5EB5CD9 vsubsd xmm3, xmm2, xmm1
    0x428A7E C5FB58C2 vaddsd xmm0, xmm0, xmm2
    0x428A82 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
    0x428A88 4D8D6D10 lea r13, [r13+0x10]
    0x428A8C C3 ret
    61 bytes, 12 instructions
    ok

    Half the size and all memory loads/stores gone.


    I went ahead and implemented float locals.
    Now I can write using JVN formula

    : Z* {f: x y u v :}
    x y f+ u f* fdup u v f+ y f* f- fswap v u f- x f* f+ ;

    this compiles to

    seea Z*
    0x42A6E0 C4C17B104D08 vmovsd xmm1, qword ptr [r13+0x8]
    0x42A6E6 C4C173584D10 vaddsd xmm1, xmm1, qword ptr [r13+0x10]
    0x42A6EC C4C173594D00 vmulsd xmm1, xmm1, qword ptr [r13]
    0x42A6F2 C4C17B585500 vaddsd xmm2, xmm0, qword ptr [r13]
    0x42A6F8 C4C16B595508 vmulsd xmm2, xmm2, qword ptr [r13+0x8]
    0x42A6FE C5F35CDA vsubsd xmm3, xmm1, xmm2
    0x42A702 C4C17B5C5500 vsubsd xmm2, xmm0, qword ptr [r13]
    0x42A708 C4C16B595510 vmulsd xmm2, xmm2, qword ptr [r13+0x10]
    0x42A70E C5EB58D1 vaddsd xmm2, xmm2, xmm1
    0x42A712 C5EB10C2 vmovsd xmm0, xmm2, xmm2
    0x42A716 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
    0x42A71C 4D8D6D10 lea r13, [r13+0x10]
    0x42A720 C3 ret
    65 bytes, 13 instructions
    4 bytes and 1 instruction more

    If I do the school book formual I get

    : Z* {f: x y u v :}
    x u f* y v f* f- x v f* y u f* f+ ;


    Which give more compact code

    seea Z*
    0x42A730 C4C17B104D00 vmovsd xmm1, qword ptr [r13]
    0x42A736 C4C173594D10 vmulsd xmm1, xmm1, qword ptr [r13+0x10]
    0x42A73C C4C17B595508 vmulsd xmm2, xmm0, qword ptr [r13+0x8]
    0x42A742 C5F35CCA vsubsd xmm1, xmm1, xmm2
    0x42A746 C4C17B595510 vmulsd xmm2, xmm0, qword ptr [r13+0x10]
    0x42A74C C4C17B105D00 vmovsd xmm3, qword ptr [r13]
    0x42A752 C4C163595D08 vmulsd xmm3, xmm3, qword ptr [r13+0x8]
    0x42A758 C5E358DA vaddsd xmm3, xmm3, xmm2
    0x42A75C C5E310C3 vmovsd xmm0, xmm3, xmm3
    0x42A760 C4C17B114D10 vmovsd qword ptr [r13+0x10], xmm1
    0x42A766 4D8D6D10 lea r13, [r13+0x10]
    0x42A76A C3 ret
    59 bytes, 12 instructions

    4 multiplications 2 add/sub

    I tried to measure the execution times. This shows no difference!

    : z1 0 do zdup zdup z* zdrop loop ;

    fpi 345e4 timer-reset 100000000 z1 .elapsed 179 ms elapsed ok
    fpi 345e4 timer-reset 100000000 z2 .elapsed 176 ms elapsed ok
    fpi 345e4 timer-reset 100000000 z3 .elapsed 177 ms elapsed ok

    I had to start with the complex number already on the stack!
    If I put fpi 345e4 inside the loop elapsed time increase with 100 times!

    BR
    Peter

    An alternative I considered was to introduce float locals. It would probably have been more complicated. The result, in the best case, would be equal in the generated code. The problem with locals is the scope, that is to the end of the word. Of course the source would be more readable!

    I might still implement float locals.

    What do you have in your system?

    BR
    Peter



    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.lang.forth on Sun Aug 30 13:31:56 2026
    From Newsgroup: comp.lang.forth

    peter <peter.noreply@tin.it> writes:
    If I do the school book formual I get

    : Z* {f: x y u v :}
    x u f* y v f* f- x v f* y u f* f+ ;


    Which give more compact code

    seea Z*
    0x42A730 C4C17B104D00 vmovsd xmm1, qword ptr [r13]
    0x42A736 C4C173594D10 vmulsd xmm1, xmm1, qword ptr [r13+0x10] >0x42A73C C4C17B595508 vmulsd xmm2, xmm0, qword ptr [r13+0x8] >0x42A742 C5F35CCA vsubsd xmm1, xmm1, xmm2
    0x42A746 C4C17B595510 vmulsd xmm2, xmm0, qword ptr [r13+0x10] >0x42A74C C4C17B105D00 vmovsd xmm3, qword ptr [r13]
    0x42A752 C4C163595D08 vmulsd xmm3, xmm3, qword ptr [r13+0x8] >0x42A758 C5E358DA vaddsd xmm3, xmm3, xmm2
    0x42A75C C5E310C3 vmovsd xmm0, xmm3, xmm3
    0x42A760 C4C17B114D10 vmovsd qword ptr [r13+0x10], xmm1
    0x42A766 4D8D6D10 lea r13, [r13+0x10]
    0x42A76A C3 ret
    59 bytes, 12 instructions

    Cool!

    I tried to measure the execution times. This shows no difference!

    Which indicates that the bottleneck of your benchmark is not in the z*
    code.

    I had to start with the complex number already on the stack!
    If I put fpi 345e4 inside the loop elapsed time increase with 100 times!

    How do you implement FP literals? One technique is to call FLIT, have
    the FP literal behind the call, and let FLIT load from its return
    address, then change its return address. That results in a branch misprediction (~20 cycles) for every FLIT, but even 2 FLITs with this
    technique do not explain 100* slowdown.

    - anton
    --
    M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
    comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
    New standard: https://forth-standard.org/
    EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From peter@peter.noreply@tin.it to comp.lang.forth on Sun Aug 30 15:58:56 2026
    From Newsgroup: comp.lang.forth

    On Sun, 30 Aug 2026 13:31:56 GMT
    anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:
    peter <peter.noreply@tin.it> writes:
    If I do the school book formual I get

    : Z* {f: x y u v :}
    x u f* y v f* f- x v f* y u f* f+ ;


    Which give more compact code

    seea Z*
    0x42A730 C4C17B104D00 vmovsd xmm1, qword ptr [r13]
    0x42A736 C4C173594D10 vmulsd xmm1, xmm1, qword ptr [r13+0x10] >0x42A73C C4C17B595508 vmulsd xmm2, xmm0, qword ptr [r13+0x8] >0x42A742 C5F35CCA vsubsd xmm1, xmm1, xmm2
    0x42A746 C4C17B595510 vmulsd xmm2, xmm0, qword ptr [r13+0x10] >0x42A74C C4C17B105D00 vmovsd xmm3, qword ptr [r13]
    0x42A752 C4C163595D08 vmulsd xmm3, xmm3, qword ptr [r13+0x8] >0x42A758 C5E358DA vaddsd xmm3, xmm3, xmm2
    0x42A75C C5E310C3 vmovsd xmm0, xmm3, xmm3
    0x42A760 C4C17B114D10 vmovsd qword ptr [r13+0x10], xmm1
    0x42A766 4D8D6D10 lea r13, [r13+0x10]
    0x42A76A C3 ret
    59 bytes, 12 instructions

    Cool!

    I tried to measure the execution times. This shows no difference!

    Which indicates that the bottleneck of your benchmark is not in the z*
    code.

    I had to start with the complex number already on the stack!
    If I put fpi 345e4 inside the loop elapsed time increase with 100 times!

    How do you implement FP literals? One technique is to call FLIT, have
    the FP literal behind the call, and let FLIT load from its return
    address, then change its return address. That results in a branch misprediction (~20 cycles) for every FLIT, but even 2 FLITs with this technique do not explain 100* slowdown.
    I found the main problem. It was FPI that was implemented as a call
    to the C math library I use. It goes thru several calls saves all regs
    and switches the stack to the original (correctly aligned).
    Now it is defined as
    $4009|21FB|5444|2D18 pad ! pad f@ FCONSTANT FPI
    This compiles to the token code
    see fpi macro
    Address OP Instruction
    0x720300 81 182D4454FB210940 FLIT 0x400921FB54442D18
    0x720309 25 RET
    10 bytes, 2 instructions
    It is a macro so will get inlined
    The assembler code becomes
    seea fpi
    0x423AD0 48B8182D4454FB210940 mov rax, 0x400921FB54442D18
    0x423ADA C4E1F96EC8 vmovq xmm1, rax
    0x423ADF C4C17B1145F8 vmovsd qword ptr [r13-0x8], xmm0
    0x423AE5 C5F310C1 vmovsd xmm0, xmm1, xmm1
    0x423AE9 4D8D6DF8 lea r13, [r13-0x8]
    0x423AED C3 ret
    30 bytes, 6 instructions
    ok
    2 instruction to load the litteral in a xmm register
    It is a bit long but I do not think it can be better
    The other possible solution would be to store litterals in memmory
    and load from there. I think this is what Aarch64 compilers do
    Peter

    - anton
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.lang.forth on Sun Aug 30 15:12:00 2026
    From Newsgroup: comp.lang.forth

    peter <peter.noreply@tin.it> writes:
    The other possible solution would be to store litterals in memmory
    and load from there. I think this is what Aarch64 compilers do

    That's also what gcc does on AMD64:

    [/tmp:171587] cat xxx.c
    double f(double x)
    {
    return x+1234.0;
    }
    [/tmp:171588] gcc -O -S xxx.c
    [/tmp:171589] cat xxx.s
    .file "xxx.c"
    .text
    .globl f
    .type f, @function
    f:
    .LFB0:
    .cfi_startproc
    addsd .LC0(%rip), %xmm0
    ret
    .cfi_endproc
    .LFE0:
    .size f, .-f
    .section .rodata.cst8,"aM",@progbits,8
    .align 8
    .LC0:
    .long 0
    .long 1083394048
    .ident "GCC: (Debian 12.2.0-14) 12.2.0"
    .section .note.GNU-stack,"",@progbits

    - anton
    --
    M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
    comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
    New standard: https://forth-standard.org/
    EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
    --- Synchronet 3.22a-Linux NewsLink 1.2