• Consequential RCU

    From Joseph Seigh@jseigh_es00@xemaps.com to comp.lang.c++ on Wed Sep 16 18:25:03 2026
    From Newsgroup: comp.lang.c++

    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-sequential-rcu/

    The name doesn't really mean anything. I was just generalizing something
    that's been around forever.* It's simple enough that it must be in use somewhere. I just haven't noticed it. I looked up the linux kernal
    seqlock documentation and it specifically says not to use pointers in
    seqlock protected data and no idea why that is unless the linux kernel
    is doing something weird with memory mapping.

    Like seqlock, it is obstruction-free. If you have qsbr (RCU) or ebr,
    that would be a better choice. It might be good for libraries where
    the memory reclamation is hidden from the application.

    * tangentially related to something that I was working on which if far
    more perverse but mostly in a solution that is looking for a problem
    phase so it probably won't go anywhere.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Joseph Seigh@jseigh_es00@xemaps.com to comp.lang.c++ on Fri Sep 18 16:07:28 2026
    From Newsgroup: comp.lang.c++

    On 9/16/26 6:25 PM, Joseph Seigh wrote:

    * tangentially related to something that I was working on which if far
    more perverse but mostly in a solution that is looking for a problem
    phase so it probably won't go anywhere.

    Actually I just though of a 2nd use for this. The POC for the first
    case was more work than I felt like doing. Ditto for the 2nd case.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Sat Sep 19 13:34:45 2026
    From Newsgroup: comp.lang.c++

    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- sequential-rcu/

    The name doesn't really mean anything. I was just generalizing something that's been around forever.*-a It's simple enough that it must be in use somewhere.-a I just haven't noticed it.-a I looked up the linux kernal seqlock documentation and it specifically says not to use pointers in
    seqlock protected data and no idea why that is unless the linux kernel
    is doing something weird with memory mapping.

    Like seqlock, it is obstruction-free.-a If you have qsbr (RCU) or ebr,
    that would be a better choice.-a It might be good for libraries where
    the memory reclamation is hidden from the application.

    * tangentially related to something that I was working on which if far
    more perverse but mostly in a solution that is looking for a problem
    phase so it probably won't go anywhere.

    Fwiw, I have used seqlocks for the writer side and pure RCU for the
    readers. Actually, there is a way to use seqlocks to gain DWCAS on a
    system that does not support it.

    https://groups.google.com/g/lock-free/c/X3fuuXknQF0/m/zfHnoFi-VXgJ

    https://pastebin.com/raw/TgTcfYtR


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Joseph Seigh@jseigh_es00@xemaps.com to comp.lang.c++ on Sun Sep 20 10:58:00 2026
    From Newsgroup: comp.lang.c++

    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
    sequential-rcu/

    The name doesn't really mean anything. I was just generalizing something
    that's been around forever.*-a It's simple enough that it must be in use
    somewhere.-a I just haven't noticed it.-a I looked up the linux kernal
    seqlock documentation and it specifically says not to use pointers in
    seqlock protected data and no idea why that is unless the linux kernel
    is doing something weird with memory mapping.

    Like seqlock, it is obstruction-free.-a If you have qsbr (RCU) or ebr,
    that would be a better choice.-a It might be good for libraries where
    the memory reclamation is hidden from the application.

    * tangentially related to something that I was working on which if far
    more perverse but mostly in a solution that is looking for a problem
    phase so it probably won't go anywhere.

    Fwiw, I have used seqlocks for the writer side and pure RCU for the
    readers. Actually, there is a way to use seqlocks to gain DWCAS on a
    system that does not support it.

    https://groups.google.com/g/lock-free/c/X3fuuXknQF0/m/zfHnoFi-VXgJ

    https://pastebin.com/raw/TgTcfYtR



    Java did its AtomicStampedReference by creating an object to hold
    the stamp and the reference and then doing a single word compare
    and swap on a reference to it. All this so you can run on
    hardware which effectively doesn't exist anymore and even if it
    did you would need to backport the os to support obsolete
    hardware so the jvm could even run on it.

    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    Re the OP, I did think of a really good POC for this but it is
    way too much work. I have better things to waste my time
    with. :)




    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Sun Sep 20 14:51:16 2026
    From Newsgroup: comp.lang.c++

    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
    sequential-rcu/

    The name doesn't really mean anything. I was just generalizing something >>> that's been around forever.*-a It's simple enough that it must be in use >>> somewhere.-a I just haven't noticed it.-a I looked up the linux kernal
    seqlock documentation and it specifically says not to use pointers in
    seqlock protected data and no idea why that is unless the linux kernel
    is doing something weird with memory mapping.

    Like seqlock, it is obstruction-free.-a If you have qsbr (RCU) or ebr,
    that would be a better choice.-a It might be good for libraries where
    the memory reclamation is hidden from the application.

    * tangentially related to something that I was working on which if far
    more perverse but mostly in a solution that is looking for a problem
    phase so it probably won't go anywhere.

    Fwiw, I have used seqlocks for the writer side and pure RCU for the
    readers. Actually, there is a way to use seqlocks to gain DWCAS on a
    system that does not support it.

    https://groups.google.com/g/lock-free/c/X3fuuXknQF0/m/zfHnoFi-VXgJ

    https://pastebin.com/raw/TgTcfYtR



    Java did its AtomicStampedReference by creating an object to hold
    the stamp and the reference and then doing a single word compare
    and swap on a reference to it.-a All this so you can run on
    hardware which effectively doesn't exist anymore and even if it
    did you would need to backport the os to support obsolete
    hardware so the jvm could even run on it.

    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    Re the OP, I did think of a really good POC for this but it is
    way too much work.-a I have better things to waste my time
    with. :)





    ;^D I will look at it, perhaps tonight. Might have some free time.
    Thanks Joe, as always: Your works is rather grand. You taught me a lot
    back in c.p.t. Thanks for that., the SenderX files? ;^o

    Fwiw, I am working on some audio for my special vector fields. Here is
    an example, A Youtube link. also, I will send you a DM over on the god
    damn FB. :^)

    https://youtu.be/gVWAgO5PAfI

    Beam me up? lol. ;^)
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Sun Sep 20 16:09:31 2026
    From Newsgroup: comp.lang.c++

    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
    sequential-rcu/

    The name doesn't really mean anything. I was just generalizing something >>> that's been around forever.*-a It's simple enough that it must be in use >>> somewhere.-a I just haven't noticed it.-a I looked up the linux kernal
    seqlock documentation and it specifically says not to use pointers in
    seqlock protected data and no idea why that is unless the linux kernel
    is doing something weird with memory mapping.

    Like seqlock, it is obstruction-free.-a If you have qsbr (RCU) or ebr,
    that would be a better choice.-a It might be good for libraries where
    the memory reclamation is hidden from the application.

    * tangentially related to something that I was working on which if far
    more perverse but mostly in a solution that is looking for a problem
    phase so it probably won't go anywhere.

    Fwiw, I have used seqlocks for the writer side and pure RCU for the
    readers. Actually, there is a way to use seqlocks to gain DWCAS on a
    system that does not support it.

    https://groups.google.com/g/lock-free/c/X3fuuXknQF0/m/zfHnoFi-VXgJ

    https://pastebin.com/raw/TgTcfYtR



    Java did its AtomicStampedReference by creating an object to hold
    the stamp and the reference and then doing a single word compare
    and swap on a reference to it.-a All this so you can run on
    hardware which effectively doesn't exist anymore and even if it
    did you would need to backport the os to support obsolete
    hardware so the jvm could even run on it.

    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    never really got into inline asm. Well. I said f it and used good ol
    GAS. Or MASM.

    GAS.

    https://web.archive.org/web/20060214112345/http://appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_gcc_asm.html

    https://web.archive.org/web/20060214112539/http://appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_masm_asm.html

    GAS was kind to me.


    # Copyright 2005 Chris Thomasson


    .align 16
    .globl np_ac_i686_atomic_dwcas_fence
    np_ac_i686_atomic_dwcas_fence:
    pushl %esi
    pushl %ebx
    movl 16(%esp), %esi
    movl (%esi), %eax
    movl 4(%esi), %edx
    movl 20(%esp), %esi
    movl (%esi), %ebx
    movl 4(%esi), %ecx
    movl 12(%esp), %esi
    lock cmpxchg8b (%esi)
    jne np_ac_i686_atomic_dwcas_fence_fail
    xorl %eax, %eax
    popl %ebx
    popl %esi
    ret

    np_ac_i686_atomic_dwcas_fence_fail:
    movl 16(%esp), %esi
    movl %eax, (%esi)
    movl %edx, 4(%esi)
    movl $1, %eax
    popl %ebx
    popl %esi
    ret




    .align 16
    .globl ac_i686_stack_mpmc_push_cas
    ac_i686_stack_mpmc_push_cas:
    movl 4(%esp), %edx
    movl (%edx), %eax
    movl 8(%esp), %ecx

    ac_i686_stack_mpmc_push_cas_retry:
    movl %eax, (%ecx)
    lock cmpxchgl %ecx, (%edx)
    jne ac_i686_stack_mpmc_push_cas_retry
    ret




    .align 16
    .globl np_ac_i686_lfgc_smr_stack_mpmc_pop_dwcas np_ac_i686_lfgc_smr_stack_mpmc_pop_dwcas:
    pushl %esi
    pushl %ebx

    np_ac_i686_lfgc_smr_stack_mpmc_pop_dwcas_reload:
    movl 12(%esp), %esi
    movl 4(%esi), %edx
    movl (%esi), %eax

    np_ac_i686_lfgc_smr_stack_mpmc_pop_dwcas_retry:
    movl 16(%esp), %ebx
    movl %eax, (%ebx)
    mfence
    cmpl (%esi), %eax
    jne np_ac_i686_lfgc_smr_stack_mpmc_pop_dwcas_reload
    test %eax, %eax
    je np_ac_i686_lfgc_smr_stack_mpmc_pop_dwcas_fail
    movl (%eax), %ebx
    leal 1(%edx), %ecx
    lock cmpxchg8b (%esi)
    jne np_ac_i686_lfgc_smr_stack_mpmc_pop_dwcas_retry

    np_ac_i686_lfgc_smr_stack_mpmc_pop_dwcas_fail:
    movl 16(%esp), %esi
    xorl %ebx, %ebx
    movl %ebx, (%esi)
    popl %ebx
    popl %esi
    ret




    .align 16
    .globl np_ac_i686_stack_mpmc_pop_dwcas
    np_ac_i686_stack_mpmc_pop_dwcas:
    pushl %esi
    pushl %ebx
    movl 12(%esp), %esi
    movl 4(%esi), %edx
    movl (%esi), %eax

    np_ac_i686_stack_mpmc_pop_dwcas_retry:
    test %eax, %eax
    je np_ac_i686_stack_mpmc_pop_dwcas_fail
    movl (%eax), %ebx
    leal 1(%edx), %ecx
    lock cmpxchg8b (%esi)
    jne np_ac_i686_stack_mpmc_pop_dwcas_retry

    np_ac_i686_stack_mpmc_pop_dwcas_fail:
    popl %ebx
    popl %esi
    ret




    .align 16
    .globl ac_i686_lfgc_smr_activate
    ac_i686_lfgc_smr_activate:
    movl 4(%esp), %edx
    movl 8(%esp), %ecx

    ac_i686_lfgc_smr_activate_reload:
    movl (%ecx), %eax
    movl %eax, (%edx)
    mfence
    cmpl (%ecx), %eax
    jne ac_i686_lfgc_smr_activate_reload
    ret




    .align 16
    .globl ac_i686_lfgc_smr_deactivate
    ac_i686_lfgc_smr_deactivate:
    movl 4(%esp), %ecx
    xorl %eax, %eax
    movl %eax, (%ecx)
    ret




    .align 16
    .globl ac_i686_queue_spsc_push
    ac_i686_queue_spsc_push:
    movl 4(%esp), %eax
    movl 8(%esp), %ecx
    movl 4(%eax), %edx
    # sfence may be needed here for future x86
    movl %ecx, (%edx)
    movl %ecx, 4(%eax)
    ret




    .align 16
    .globl ac_i686_queue_spsc_pop
    ac_i686_queue_spsc_pop:
    pushl %ebx
    movl 8(%esp), %ecx
    movl (%ecx), %eax
    cmpl 4(%ecx), %eax
    je ac_i686_queue_spsc_pop_failed
    movl (%eax), %edx
    # lfence may be needed here for future x86
    movl 12(%edx), %ebx
    movl %edx, (%ecx)
    movl %ebx, 12(%eax)
    popl %ebx
    ret

    ac_i686_queue_spsc_pop_failed:
    xorl %eax, %eax
    popl %ebx
    ret




    .align 16
    .globl ac_i686_mb_fence
    ac_i686_mb_fence:
    mfence
    ret




    .align 16
    .globl ac_i686_mb_naked
    ac_i686_mb_naked:
    ret




    .align 16
    .globl ac_i686_mb_store_fence
    ac_i686_mb_store_fence:
    movl 4(%esp), %ecx
    movl 8(%esp), %eax
    mfence
    movl %eax, (%ecx)
    ret




    .align 16
    .globl ac_i686_mb_store_naked
    ac_i686_mb_store_naked:
    movl 4(%esp), %ecx
    movl 8(%esp), %eax
    movl %eax, (%ecx)
    ret




    .align 16
    .globl ac_i686_mb_load_fence
    ac_i686_mb_load_fence:
    movl 4(%esp), %ecx
    movl (%ecx), %eax
    mfence
    ret




    .align 16
    .globl ac_i686_mb_load_naked
    ac_i686_mb_load_naked:
    movl 4(%esp), %ecx
    movl (%ecx), %eax
    ret




    .align 16
    .globl ac_i686_atomic_xchg_fence
    ac_i686_atomic_xchg_fence:
    movl 4(%esp), %ecx
    movl 8(%esp), %eax
    xchgl %eax, (%ecx)
    ret




    .align 16
    .globl ac_i686_atomic_xadd_fence
    ac_i686_atomic_xadd_fence:
    movl 4(%esp), %ecx
    movl 8(%esp), %eax
    lock xaddl %eax, (%ecx)
    ret




    .align 16
    .globl ac_i686_atomic_inc_fence
    ac_i686_atomic_inc_fence:
    movl 4(%esp), %ecx
    movl $1, %eax
    lock xaddl %eax, (%ecx)
    incl %eax
    ret




    .align 16
    .globl ac_i686_atomic_dec_fence
    ac_i686_atomic_dec_fence:
    movl 4(%esp), %ecx
    movl $-1, %eax
    lock xaddl %eax, (%ecx)
    decl %eax
    ret




    .align 16
    .globl ac_i686_atomic_cas_fence
    ac_i686_atomic_cas_fence:
    movl 4(%esp), %ecx
    movl 8(%esp), %eax
    movl 12(%esp), %edx
    lock cmpxchgl %edx, (%ecx)
    ret





    Re the OP, I did think of a really good POC for this but it is
    way too much work.-a I have better things to waste my time
    with. :)





    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Joseph Seigh@jseigh_es00@xemaps.com to comp.lang.c++ on Sun Sep 20 20:19:04 2026
    From Newsgroup: comp.lang.c++

    On 9/20/26 7:09 PM, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:

    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -anever really got into inline asm. Well. I said f it and used good ol
    GAS. Or MASM.

    GAS.

    https://web.archive.org/web/20060214112345/http:// appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_gcc_asm.html

    https://web.archive.org/web/20060214112539/http:// appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_masm_asm.html

    GAS was kind to me.


    https://github.com/jseigh/queues/blob/main/include/atomix.h
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Sun Sep 20 19:26:39 2026
    From Newsgroup: comp.lang.c++

    On 9/20/2026 5:19 PM, Joseph Seigh wrote:
    On 9/20/26 7:09 PM, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:

    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -a-anever really got into inline asm. Well. I said f it and used good ol
    GAS. Or MASM.

    GAS.

    https://web.archive.org/web/20060214112345/http://
    appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_gcc_asm.html

    https://web.archive.org/web/20060214112539/http://
    appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_masm_asm.html

    GAS was kind to me.


    https://github.com/jseigh/queues/blob/main/include/atomix.h

    Excellent! 64-bit. Notice my DWCAS. Looks familiar, right? I basically
    ported yours way back in the day. I think I first saw yours in c.p.t, or
    that asm usenet group you also posted in.

    Returning 0 for success, was just POSIX habits personification 101? ;^D

    In MASM:

    align 16
    np_ac_i686_atomic_dwcas_fence PROC
    push esi
    push ebx
    mov esi, [esp + 16]
    mov eax, [esi]
    mov edx, [esi + 4]
    mov esi, [esp + 20]
    mov ebx, [esi]
    mov ecx, [esi + 4]
    mov esi, [esp + 12]
    lock cmpxchg8b qword ptr [esi]
    jne np_ac_i686_atomic_dwcas_fence_fail
    xor eax, eax
    pop ebx
    pop esi
    ret

    np_ac_i686_atomic_dwcas_fence_fail:
    mov esi, [esp + 16]
    mov [esi + 0], eax;
    mov [esi + 4], edx;
    mov eax, 1
    pop ebx
    pop esi
    ret
    np_ac_i686_atomic_dwcas_fence ENDP
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Mon Sep 21 14:50:19 2026
    From Newsgroup: comp.lang.c++

    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
    sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -anever really got into inline asm. Well. I said f it and used good ol
    GAS. Or MASM.


    You really should learn about inline assembly in gcc. It can be
    significantly more efficient than using an external call in a separately assembled file. And I am confident that you'd find it fun! (Of course,
    it means the code is somewhat non-portable, but portability has never
    been a strength of assembly. The same syntax is supported by clang,
    icc, and some other compilers.)

    To be clear here - I have no comment on the actual use of the
    instructions or the locking algorithms. You and Joseph know these much
    better than I do, especially in the x86 world. I am merely showing what
    I think is a better way to write the code and integrate it with C++, at
    least for gcc and other compilers supporting the same syntax. I am also
    not sure if this is a "strong" or "weak" compare-exchange, or the exact
    memory ordering. That is not really important for showing the inline assembly.

    Your intention (AFAICS) here is to wrap the locked cmpxchg8b and
    cmpxchg16b instructions. The 8-byte version takes a uint64_t * pointer
    "p" and compares it to EDX:EAX. If it is equal, ZF is set and *p is set
    to ECX:EBX. Otherwise ZF is clearer and EDX:EAX is loaded with *p. The 16-byte version is similar, but with RDX:RAX and RCX:RBX. There is also
    the cmpxchg instruction that works takes a T* pointer p and a register
    R, and uses AL/AX/EAX/RAX and register R instead of register pairs. T
    can be up to 32-bit in x86-32, and 64-bit in x86-64.

    Since x86-32 is pretty much outdated, let's stick to x86-64 and make two functions. However, gains from good inline assembly can be higher in
    32-bit mode as it can save a great deal of register/stack manipulation
    dues to the terrible x86-32 ABI for function calls.

    // Code here works for C and C++, though calling this "uint128_t"
    // is cheating a bit.

    typedef unsigned __int128 uint128_t;

    // Semantics - atomically do :
    // if (*p_var == *p_expected) {
    / *p_var = desired;
    // return true;
    // } else {
    / *p_expected = *p_var;
    // return false;
    // }


    bool compare_exchange8(volatile uint64_t * p_var,
    uint64_t * p_expected, uint64_t desired);

    bool compare_exchange16(volatile uint128_t * p_var,
    uint128_t * p_expected, uint128_t desired);


    The 8-byte case is easy - use C++'s standard atomic compare_exchange(),
    or for C11 onwards, use standard atomic_compare_exchange_strong(). For pre-C11/C++11, use gcc's __atomic_compare_exchange_n(). Testing with
    godbolt gives the same code in each case :

    bool cpp_compare_exchange8(volatile std::atomic<uint64_t> * p_var,
    uint64_t * p_expected, uint64_t desired) {

    return p_var->compare_exchange_strong(*p_expected, desired);
    }

    "cpp_compare_exchange8(std::atomic<unsigned long> volatile*, unsigned
    long*, unsigned long)":
    mov rax, QWORD PTR [rsi]
    lock cmpxchg QWORD PTR [rdi], rdx
    je .L8
    mov QWORD PTR [rsi], rax
    .L8:
    sete al
    ret

    Although you clearly should be using the standard library functions
    rather than home-made inline assembly, let's do this with inline
    assembly anyway as an example :

    inline bool inline_compare_exchange8(volatile uint64_t * p_var,
    uint64_t * p_expected, uint64_t desired) {

    bool result;
    uint64_t expected = *p_expected;
    asm volatile(
    "lock cmpxchgq %[desired], (%[p_var])"
    : [result] "=@ccz" (result),
    [expected] "+a" (expected)
    : [p_var] "r" (p_var),
    [desired] "r" (desired)
    : "memory", "cc"
    );
    if (!result) *p_expected = expected;
    return result;
    }

    Key points here include :

    1. Use a "bool" result that matches the "=@ccz" constraint to capture
    the zero flag directly - code that calls this function can immediately
    use "je" or "cmove" instructions, because the compiler knows the result
    is the zero flag.

    2. "expected" is an input/output operand.

    3. Registers are allocated automatically by the compiler, except for "expected" which must use the "A" register (as that's how the "cmpxchgq" instruction works). This means that if calling code already has the
    values in other registers, it uses them directly - no extra register
    moves or stacking.

    4. The operands to the assembly are given explicit names so that the
    code is clearer. (They don't have to match the C names, but it's
    usually easiest if they do.)

    5. Local C variables can be made and used freely - they do not lead to
    any extra costs in the generated code.

    6. The inline assembly itself is absolutely minimal - a single
    instruction. That should always be the aim for inline assembly. The
    more the compiler does itself, the less you have to do manually, and the
    more efficient the end results. It's a win-win situation.


    The generated code here is identical to that of the standard library
    version. In the link below, I also have a variant for -m32 code, and a
    quick test function that shows that the zero flag works as desired.
    There are a couple of additional register moves in the test of the
    32-bit version (compared to the standard library version), but otherwise
    the code is identical and, AFAICS, optimal.

    <https://godbolt.org/z/hEcPzhGr5>


    The 16-byte version is a bit more involved, because the 128-bit
    "expected" and "desired" need to be split up and put in specific
    registers. (Just like the 32-bit version of the 8-byte function.) But
    the same principles apply - do all that manipulation in C++, not
    assembly. This also makes it much easier to adapt to a real function in
    which you might not have 128-bit "expected" and "desired", but four
    separate 64-bit items (values, addresses, etc.).

    inline bool inline_compare_exchange16(volatile uint128_t * p_var,
    uint128_t * p_expected, uint128_t desired) {

    bool result;
    uint128_t expected = *p_expected;
    uint64_t expected_lo = (uint64_t) expected;
    uint64_t expected_hi = (uint64_t) (expected >> 64);
    const uint64_t desired_lo = (uint64_t) desired;
    const uint64_t desired_hi = (uint64_t) (desired >> 64);
    asm volatile(
    "lock cmpxchg16b (%[p_var])"
    : [result] "=@ccz" (result),
    [expected_lo] "+a" (expected_lo),
    [expected_hi] "+d" (expected_hi)
    : [p_var] "r" (p_var),
    [desired_lo] "b" (desired_lo),
    [desired_hi] "c" (desired_hi)
    : "memory", "cc"
    );
    if (!result) {
    expected = ((uint128_t) expected_hi << 64) + expected_lo;
    *p_expected = expected;
    }
    return result;
    }

    Again, there is only one line of assembly here. The generated code is:

    inline_compare_exchange16(unsigned __int128 volatile*, unsigned
    __int128*, unsigned __int128):
    pushq %rbx
    movq (%rsi), %rax
    movq %rdx, %rbx
    movq 8(%rsi), %rdx
    lock cmpxchg16b (%rdi)
    je .L12
    movq %rax, (%rsi)
    movq %rdx, 8(%rsi)
    .L12:
    sete %al
    popq %rbx
    ret

    This is a bit more efficient than Joseph's version. The gains will be
    more noticeable in code that uses the function as there will be fewer
    register moves, no need to push variables out of registers and onto the
    stack before loading them into registers again, and the zero flag can be
    used directly.


    David






    https://web.archive.org/web/20060214112345/http:// appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_gcc_asm.html

    https://web.archive.org/web/20060214112539/http:// appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_masm_asm.html

    GAS was kind to me.


    # Copyright 2005 Chris Thomasson


    .align 16
    .globl np_ac_i686_atomic_dwcas_fence
    np_ac_i686_atomic_dwcas_fence:
    -a pushl %esi
    -a pushl %ebx
    -a movl 16(%esp), %esi
    -a movl (%esi), %eax
    -a movl 4(%esi), %edx
    -a movl 20(%esp), %esi
    -a movl (%esi), %ebx
    -a movl 4(%esi), %ecx
    -a movl 12(%esp), %esi
    -a lock cmpxchg8b (%esi)
    -a jne np_ac_i686_atomic_dwcas_fence_fail
    -a xorl %eax, %eax
    -a popl %ebx
    -a popl %esi
    -a ret

    np_ac_i686_atomic_dwcas_fence_fail:
    -a movl 16(%esp), %esi
    -a movl %eax, (%esi)
    -a movl %edx, 4(%esi)
    -a movl $1, %eax
    -a popl %ebx
    -a popl %esi
    ret




    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From boltar@boltar@caprica.universe to comp.lang.c++ on Mon Sep 21 15:47:10 2026
    From Newsgroup: comp.lang.c++

    On Mon, 21 Sep 2026 14:50:19 +0200
    David Brown <david.brown@hesbynett.no> gabbled:
    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
    sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -anever really got into inline asm. Well. I said f it and used good ol
    GAS. Or MASM.


    You really should learn about inline assembly in gcc. It can be >significantly more efficient than using an external call in a separately >assembled file. And I am confident that you'd find it fun! (Of course,

    Unless the assembler is just a few opcodes you're better off leaving it to
    a compiler as it'll usually do a better job and probably knows opcodes you've never even heard of.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Mon Sep 21 18:45:19 2026
    From Newsgroup: comp.lang.c++

    On 21/09/2026 17:47, boltar@caprica.universe wrote:
    On Mon, 21 Sep 2026 14:50:19 +0200
    David Brown <david.brown@hesbynett.no> gabbled:
    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- >>>>>> sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -a-anever really got into inline asm. Well. I said f it and used good
    ol GAS. Or MASM.


    You really should learn about inline assembly in gcc.-a It can be
    significantly more efficient than using an external call in a
    separately assembled file.-a And I am confident that you'd find it
    fun!-a (Of course,

    Unless the assembler is just a few opcodes you're better off leaving it to
    a compiler as it'll usually do a better job and probably knows opcodes you've
    never even heard of.


    The point here was an opcode - the 16 byte compare-and-exchange - that
    the compiler does /not/ know about, and is not supported by the C++ (or
    C) standard library for atomics. And the point of the inline assembly I
    gave is that it is just one line of actual assembly, precisely because
    the compiler does a better job of all the "housekeeping" stuff of moving things in and out of registers than you can do with manual assembly code.

    Next time, perhaps you might like to read more than the first two lines
    of a post before replying to it.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Thu Sep 24 15:28:32 2026
    From Newsgroup: comp.lang.c++

    On 9/21/2026 5:50 AM, David Brown wrote:
    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
    sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -a-anever really got into inline asm. Well. I said f it and used good ol
    GAS. Or MASM.


    You really should learn about inline assembly in gcc.-a It can be significantly more efficient than using an external call in a separately assembled file.-a And I am confident that you'd find it fun!-a (Of course, [...]

    Yeah. Well. Shit. Back then I did not want any compiler to mess with
    things, and my inline skills were rather lacking! I felt way better
    making a header the structs, and externally assembling things. Keep in
    mind this was before C/C++ 11. I even said woe to LTO.

    Iirc this was my abstraction for C over it:

    https://web.archive.org/web/20060214112519/http://appcore.home.comcast.net/appcore/include/cpu/i686/ac_i686_h.html

    Just a simple layer for C to connect to my ASM:

    /* Copyright 2005 Chris Thomasson */


    #ifndef AC_I686_H
    #define AC_I686_H


    #ifdef __cplusplus
    extern "C"
    {
    #endif




    /***** i686 Specific API *****/
    #define AC_I686_CACHE_LINE 128
    #define AC_I686_WORD_SIZE 4
    #define AC_DECLSPEC_ALIGN_CACHE_LINE AC_DECLSPEC_ALIGN( AC_I686_CACHE_LINE )


    #ifdef _MSC_VER
    typedef signed __int32 ac_i686_intword_t;
    typedef unsigned __int32 ac_i686_uintword_t;
    typedef signed __int16 ac_i686_intword2_t;
    typedef unsigned __int16 ac_i686_uintword2_t;
    #define ac_i686_pause() _asm pause

    #else
    typedef signed int ac_i686_intword_t;
    typedef unsigned int ac_i686_uintword_t;
    typedef signed int short ac_i686_intword2_t;
    typedef unsigned int short ac_i686_uintword2_t;
    #define ac_i686_pause() __asm__ __volatile__ ( "pause" )
    #endif


    typedef ac_i686_uintword_t ac_i686_flags_t;


    #define ac_i686_pause_yield() ac_i686_pause(); sched_yield()




    #ifdef _MSC_VER
    /* 4324: structure was padded due to __declspec(align()) */
    #pragma warning ( disable : 4324 )
    #pragma pack(1)
    #endif

    /* MUST be four adjacent words */
    typedef struct
    AC_DECLSPEC_PACKED
    ac_i686_node_
    {
    struct ac_i686_node_ *next;
    struct ac_i686_node_ *lfgc_next;
    ac_fp_dtor_t fp_dtor;
    const void *state;

    } ac_i686_node_t;


    /* MUST be two adjacent words. front must be at offset 0 */
    typedef struct
    AC_DECLSPEC_PACKED
    ac_i686_stack_mpmc_
    {
    ac_i686_node_t *front;
    ac_i686_uintword_t aba;

    } ac_i686_stack_mpmc_t;


    /* MUST be two adjacent words. front must be at offset 0 */
    typedef struct
    AC_DECLSPEC_PACKED
    ac_i686_queue_spsc_
    {
    ac_i686_node_t *front;
    ac_i686_node_t *back;

    } ac_i686_queue_spsc_t;


    /* MUST be two adjacent words */
    typedef struct
    AC_DECLSPEC_PACKED
    ac_i686_tls_
    {
    struct ac_i686_lfgc_smr_ *lfgc_smr;
    ac_thread_t *thread;

    } ac_i686_tls_t;


    #ifdef _MSC_VER
    #pragma pack()
    /* 4324: structure was padded due to __declspec(align()) */
    #pragma warning ( default : 4324 )
    #endif




    /* critical i686 atomic api compile time assertion */
    AC_BUILD_DBG_ASSERT
    ( i686,
    sizeof( void* ) == AC_I686_WORD_SIZE &&
    sizeof( ac_fp_dtor_t ) == AC_I686_WORD_SIZE &&
    sizeof( ac_i686_node_t* ) == AC_I686_WORD_SIZE &&
    sizeof( ac_i686_tls_t* ) == AC_I686_WORD_SIZE &&
    sizeof( ac_i686_queue_spsc_t* ) == AC_I686_WORD_SIZE &&
    sizeof( struct ac_i686_lfgc_smr_* ) == AC_I686_WORD_SIZE &&
    sizeof( ac_thread_t* ) == AC_I686_WORD_SIZE &&
    sizeof( ac_i686_intword_t ) == AC_I686_WORD_SIZE &&
    sizeof( ac_i686_uintword_t ) == AC_I686_WORD_SIZE &&
    sizeof( ac_i686_intword2_t ) == AC_I686_WORD_SIZE / 2 &&
    sizeof( ac_i686_uintword2_t ) == AC_I686_WORD_SIZE / 2 &&
    sizeof( ac_i686_node_t ) == AC_I686_WORD_SIZE * 4 &&
    sizeof( ac_i686_queue_spsc_t ) == AC_I686_WORD_SIZE * 2 &&
    sizeof( ac_i686_tls_t ) == AC_I686_WORD_SIZE * 2 &&
    sizeof( ptrdiff_t ) == AC_I686_WORD_SIZE &&
    sizeof( size_t ) == AC_I686_WORD_SIZE );




    AC_SYS_APIEXPORT
    void AC_CDECL
    ac_i686_mb_fence
    ( void );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_mb_load_fence
    ( ac_i686_intword_t* );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_mb_store_fence
    ( ac_i686_intword_t*,
    ac_i686_intword_t );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_mb_load_naked
    ( ac_i686_intword_t* );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_mb_store_naked
    ( ac_i686_intword_t*,
    ac_i686_intword_t );


    AC_SYS_APIEXPORT
    int AC_CDECL
    np_ac_i686_atomic_dwcas_fence
    ( void*,
    void*,
    const void* );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_atomic_cas_fence
    ( ac_i686_intword_t*,
    ac_i686_intword_t,
    ac_i686_intword_t );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_atomic_xchg_fence
    ( ac_i686_intword_t*,
    ac_i686_intword_t );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_atomic_xadd_fence
    ( ac_i686_intword_t*,
    ac_i686_intword_t );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_atomic_inc_fence
    ( ac_i686_intword_t* );


    AC_SYS_APIEXPORT
    ac_i686_intword_t AC_CDECL
    ac_i686_atomic_dec_fence
    ( ac_i686_intword_t* );


    AC_SYS_APIEXPORT void AC_CDECL
    ac_i686_stack_mpmc_push_cas
    ( ac_i686_stack_mpmc_t*,
    ac_i686_node_t* );


    AC_SYS_APIEXPORT ac_i686_node_t* AC_CDECL
    np_ac_i686_stack_mpmc_pop_dwcas
    ( ac_i686_stack_mpmc_t* );


    AC_SYS_APIEXPORT void AC_CDECL
    ac_i686_queue_spsc_push
    ( ac_i686_queue_spsc_t*,
    ac_i686_node_t* );


    AC_SYS_APIEXPORT ac_i686_node_t* AC_CDECL
    ac_i686_queue_spsc_pop
    ( ac_i686_queue_spsc_t* );


    AC_APIEXPORT ac_i686_node_t* AC_APIDECL
    ac_i686_node_cache_pop
    ( const void *state );


    AC_APIEXPORT void AC_APIDECL
    ac_i686_node_cache_push
    ( ac_i686_node_t* );


    AC_APIEXPORT int AC_APIDECL
    ac_i686_node_cache_push_no_free
    ( ac_i686_node_t* );




    #define ac_i686_stack_mpmc_init( ac_macro_this ) \
    (ac_macro_this)->front = 0


    #define ac_i686_queue_spsc_init( ac_macro_this, ac_macro_dummy ) \
    (ac_macro_this)->front = (ac_macro_dummy); \
    (ac_macro_this)->back = (ac_macro_dummy)


    #define ac_i686_node_init( ac_macro_this, ac_macro_fp_dtor,
    ac_macro_state ) \
    (ac_macro_this)->next = 0; \
    (ac_macro_this)->lfgc_next = 0; \
    (ac_macro_this)->fp_dtor = (ac_macro_fp_dtor); \
    (ac_macro_this)->state = (ac_macro_state)


    #define ac_i686_node_get_next( ac_macro_this ) \
    ( (ac_macro_this)->next )


    #define ac_i686_mb_node_get_next( ac_macro_this ) \
    ( (ac_i686_node_t*)ac_mb_loadptr_depends( &(ac_macro_this)->next ) )


    #define ac_i686_node_get_state( ac_macro_this ) \
    ( (void*)(ac_macro_this)->state )


    #define ac_i686_mb_node_get_state( ac_macro_this ) \
    ( ac_mb_loadptr_depends( &(ac_macro_this)->state ) )


    #define ac_i686_mb_loadptr_fence( ac_macro_state ) \
    ( (void*)ac_i686_mb_load_fence \
    ( (ac_i686_intword_t*)(ac_macro_state) ) )


    #define ac_i686_mb_loadptr_naked( ac_macro_state ) \
    ( (void*)ac_i686_mb_load_naked \
    ( (ac_i686_intword_t*)(ac_macro_state) ) )


    #define ac_i686_mb_storeptr_fence( ac_macro_dest, ac_macro_state ) \
    ( (void*)ac_i686_mb_store_fence \
    ( (ac_i686_intword_t*)(ac_macro_dest), \
    (ac_i686_intword_t)(ac_macro_state) ) )


    #define ac_i686_mb_storeptr_naked( ac_macro_dest, ac_macro_state ) \
    ( (void*)ac_i686_mb_store_naked \
    ( (ac_i686_intword_t*)(ac_macro_dest), \
    (ac_i686_intword_t)(ac_macro_state) ) )


    #define ac_i686_atomic_casptr_fence( ac_macro_dest, ac_macro_cmp, ac_macro_xchg ) \
    ( (void*)ac_i686_atomic_cas_fence \
    ( (ac_i686_intword_t*)(ac_macro_dest), \
    (ac_i686_intword_t)(ac_macro_cmp), \
    (ac_i686_intword_t)(ac_macro_xchg) ) )


    #define ac_i686_atomic_xchgptr_fence( ac_macro_dest, ac_macro_xchg ) \
    ( (void*)ac_i686_atomic_xchg_fence \
    ( (ac_i686_intword_t*)(ac_macro_dest), \
    (ac_i686_intword_t)(ac_macro_xchg) ) )








    /***** Low-Level Atomic API Abstraction *****/
    #define AC_CPU_CACHE_LINE AC_I686_CACHE_LINE
    #define AC_CPU_WORD_SIZE AC_I686_WORD_SIZE


    #ifndef AC_CPU_HAS_DWCAS
    #define AC_CPU_HAS_DWCAS
    #endif




    typedef ac_i686_intword_t ac_intword_t;
    typedef ac_i686_uintword_t ac_uintword_t;
    typedef ac_i686_intword2_t ac_intword2_t;
    typedef ac_i686_uintword2_t ac_uintword2_t;
    typedef ac_i686_flags_t ac_flags_t;


    typedef ac_i686_node_t ac_cpu_node_t;
    typedef ac_i686_stack_mpmc_t ac_cpu_stack_mpmc_t;
    typedef ac_i686_queue_spsc_t ac_cpu_queue_spsc_t;
    typedef ac_i686_tls_t ac_cpu_tls_t;




    #define ac_cpu_pause ac_i686_pause
    #define ac_cpu_pause_yield ac_i686_pause_yield


    #define ac_mb_fence ac_i686_mb_fence


    #define ac_mb_load_fence ac_i686_mb_load_fence
    #define ac_mb_load_naked ac_i686_mb_load_naked
    #define ac_mb_load_acquire ac_mb_load_naked
    #define ac_mb_load_depends ac_mb_load_naked
    #define ac_mb_loadptr_fence ac_i686_mb_loadptr_fence
    #define ac_mb_loadptr_naked ac_i686_mb_loadptr_naked
    #define ac_mb_loadptr_acquire ac_mb_loadptr_naked
    #define ac_mb_loadptr_depends ac_mb_loadptr_naked


    #define ac_mb_store_fence ac_i686_mb_store_fence
    #define ac_mb_store_naked ac_i686_mb_store_naked
    #define ac_mb_store_release ac_mb_store_naked
    #define ac_mb_storeptr_fence ac_i686_mb_storeptr_fence
    #define ac_mb_storeptr_naked ac_i686_mb_storeptr_naked
    #define ac_mb_storeptr_release ac_mb_storeptr_naked


    #define np_ac_atomic_dwcas_fence np_ac_i686_atomic_dwcas_fence
    #define np_ac_atomic_dwcas_acquire np_ac_atomic_dwcas_fence
    #define np_ac_atomic_dwcas_release np_ac_atomic_dwcas_fence
    #define np_ac_atomic_dwcas_depends np_ac_atomic_dwcas_fence


    #define ac_atomic_cas_fence ac_i686_atomic_cas_fence
    #define ac_atomic_cas_acquire ac_atomic_cas_fence
    #define ac_atomic_cas_release ac_atomic_cas_fence
    #define ac_atomic_cas_depends ac_atomic_cas_fence
    #define ac_atomic_casptr_fence ac_i686_atomic_casptr_fence
    #define ac_atomic_casptr_acquire ac_atomic_casptr_fence
    #define ac_atomic_casptr_release ac_atomic_casptr_fence
    #define ac_atomic_casptr_depends ac_atomic_casptr_fence


    #define ac_atomic_xadd_fence ac_i686_atomic_xadd_fence
    #define ac_atomic_xadd_acquire ac_atomic_xadd_fence
    #define ac_atomic_xadd_release ac_atomic_xadd_fence
    #define ac_atomic_xadd_depends ac_atomic_xadd_fence


    #define ac_atomic_xchg_fence ac_i686_atomic_xchg_fence
    #define ac_atomic_xchg_acquire ac_atomic_xchg_fence
    #define ac_atomic_xchg_release ac_atomic_xchg_fence
    #define ac_atomic_xchg_depends ac_atomic_xchg_fence
    #define ac_atomic_xchgptr_fence ac_i686_atomic_xchgptr_fence
    #define ac_atomic_xchgptr_acquire ac_atomic_xchgptr_fence
    #define ac_atomic_xchgptr_release ac_atomic_xchgptr_fence
    #define ac_atomic_xchgptr_depends ac_atomic_xchgptr_fence


    #define ac_atomic_inc_fence ac_i686_atomic_inc_fence
    #define ac_atomic_inc_acquire ac_atomic_inc_fence
    #define ac_atomic_inc_release ac_atomic_inc_fence
    #define ac_atomic_inc_depends ac_atomic_inc_fence


    #define ac_atomic_dec_fence ac_i686_atomic_dec_fence
    #define ac_atomic_dec_acquire ac_atomic_dec_fence
    #define ac_atomic_dec_release ac_atomic_dec_fence
    #define ac_atomic_dec_depends ac_atomic_dec_fence


    #define ac_cpu_queue_spsc_init ac_i686_queue_spsc_init
    #define ac_cpu_queue_spsc_push ac_i686_queue_spsc_push
    #define ac_cpu_queue_spsc_pop ac_i686_queue_spsc_pop


    #define ac_cpu_stack_mpmc_init ac_i686_stack_mpmc_init
    #define ac_cpu_stack_mpmc_push_cas ac_i686_stack_mpmc_push_cas
    #define np_ac_cpu_stack_mpmc_pop_dwcas np_ac_i686_stack_mpmc_pop_dwcas


    #define ac_cpu_node_init ac_i686_node_init
    #define ac_cpu_node_get_next ac_i686_node_get_next
    #define ac_cpu_mb_node_get_next ac_i686_mb_node_get_next
    #define ac_cpu_node_get_state ac_i686_node_get_state
    #define ac_cpu_mb_node_get_state ac_i686_mb_node_get_state
    #define ac_cpu_node_cache_pop ac_i686_node_cache_pop
    #define ac_cpu_node_cache_push ac_i686_node_cache_push
    #define ac_cpu_node_cache_push_no_free ac_i686_node_cache_push_no_free





    #ifdef __cplusplus
    }
    #endif


    #endif




    #include "ac_i686_lfgc_smr.h"
    #include "ac_i686_lfgc_refcount.h"

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Thu Sep 24 15:31:46 2026
    From Newsgroup: comp.lang.c++

    On 9/21/2026 8:47 AM, boltar@caprica.universe wrote:
    On Mon, 21 Sep 2026 14:50:19 +0200
    David Brown <david.brown@hesbynett.no> gabbled:
    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- >>>>>> sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -a-anever really got into inline asm. Well. I said f it and used good
    ol GAS. Or MASM.


    You really should learn about inline assembly in gcc.-a It can be
    significantly more efficient than using an external call in a
    separately assembled file.-a And I am confident that you'd find it
    fun!-a (Of course,

    Unless the assembler is just a few opcodes you're better off leaving it to
    a compiler as it'll usually do a better job and probably knows opcodes you've
    never even heard of.





    Back then, I was really worried because the C/C++ stds did not give a
    shit about atomics. Well, those that claimed POSIX std support aside for
    a moment:

    https://web.archive.org/web/20071008213108/http://appcore.home.comcast.net/

    AppCore: A Portable High-Performance Thread Synchronization Library

    An Effective Marriage between Lock-Free and Lock-Based Algorithms


    This page presents a scaleable single-producer/consumer lock-free queue
    and an atomic operations api for the i686. There is also a full-blown
    hazard pointer implementation that is used to create an atomic reference
    count api. This makes it possible to create a single-word atomic C++
    smart pointer. A very crude smart pointer class is included as a proof
    of concept. The code is made up of bits and pieces from my rCLnew and improvedrCY AppCore Library. I can't reveal the entire library until I completely finish its documentation. However, the presented algorithm
    provides a highly-efficient method for thread-to-thread communication.
    The code compiles with gcc or msvc++ and relies on pthreads; no
    dependence on windows headers. It should compile with gcc on many
    operating systems that provide pthreads under i686 cpu's. The lock-free
    queue does not rely on any rCLatomic operationsrCY. It uses simple loads and stores guarded with memory barriers. All of its rCLcritical-sequencesrCY are contained in externally assembled functions ( read all ) in order to
    prevent a rouge C compiler from reordering anything that would corrupt
    the data-structure. The queue allocates its nodes from a three-level
    cache. The first level is a local per-thread LIFO, the second is a
    global lock-free LIFO, and the third is malloc/free. The lock-free data-structures are padded and forcefully aligned on separate cache
    lines. Unfortunately, all of this is necessary to achieve scaleable thread-to-thread communication. Test the presented code, and compare and contrast its performance against a "traditional lock-based" solution for
    a single-producer/consumer queue. You just may find that lock-free thread-to-thread communication can be stable, simple, and "useful" after all... ;)


    ****This site is under construction****



    Contact
    Remove nospam and underscores: nospam_cristom@nospam_comcast.net
    I also post and discuss information on AppCore to comp.programming.threads


    Links
    A Fast-Pathed POSIX Thread library

    "Mostly" Lock-Free Word-Based Atomic Reference Counting (source)


    Code Status
    - Completed: Feb. 21, 2005: Added atomic api
    - Completed: Feb. 21, 2005: Critical update! Slays simple mem leak wrt
    tls caused by stupid cut&paste; error! ;(...
    - Completed: Mar. 14, 2005: Added hazard-pointers
    - Completed: Mar. 25, 2005: Critical update! Added updated assembly
    files. Re-download
    - Completed: Mar. 30, 2005: Added simple MSVC 6.0 workspaces and Dev-C++ project files
    - Completed: Apr. 18, 2005: Fixed pthread related memory leak that only
    shows up on Linux. Re-download
    - Completed: May 2, 2005: Minor code updates and other various
    performance enhancements. Re-download
    - Completed: May 3, 2005: Found and removed a small bug wrt SMR & node
    cache. ItrCOs hard to trip, but it will crash if it does. Its fixed so you should probably re-download.
    - Completed: May 10, 2005: Found and removed the last SMR & node cache
    related bug. Again, it can be hard to trip. Re-download
    - Completed: May 14, 2005: Found and removed alloca scoping bug.
    - Currently: Verifying that AppCore is now bug free! :)


    Downloads
    appcore version: 0.0.1 / arch: i686 ( pre-alpha )
    The MSVC 6.0 workspaces and the Dev-C++ projects build the appcore.dll
    in the c:/winnt/system32 directory. This may need to be changed to suite
    your needs.


    Project File Index(s)
    appcore project
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Fri Sep 25 08:33:44 2026
    From Newsgroup: comp.lang.c++

    On 25/09/2026 00:28, Chris M. Thomasson wrote:
    On 9/21/2026 5:50 AM, David Brown wrote:
    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- >>>>>> sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -a-anever really got into inline asm. Well. I said f it and used good
    ol GAS. Or MASM.


    You really should learn about inline assembly in gcc.-a It can be
    significantly more efficient than using an external call in a
    separately assembled file.-a And I am confident that you'd find it
    fun!-a (Of course, [...]

    Yeah. Well. Shit. Back then I did not want any compiler to mess with
    things, and my inline skills were rather lacking! I felt way better
    making a header the structs, and externally assembling things. Keep in
    mind this was before C/C++ 11. I even said woe to LTO.


    I was not trying to show pre-C++11 Chris how to write inline assembly
    for atomics. I was trying to show 2026 Chris how to write inline
    assembly /today/, because it is still useful for some things (like a
    16-byte compare-and-swap). I was also trying to show Joseph how to do a better job than he had so far. (Again, it's on the inline assembly
    part, not the choice of instructions or atomic handling.)

    You didn't know the details of gcc inline assembly back then - fair
    enough, it's not easy, and has subtly challenges. Now you know a bit
    more going forward. And I really do think it is something that you
    would enjoy playing with.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Bonita Montero@Bonita.Montero@gmail.com to comp.lang.c++ on Fri Sep 25 18:25:41 2026
    From Newsgroup: comp.lang.c++

    Am 21.09.2026 um 18:45 schrieb David Brown:

    The point here was an opcode - the 16 byte compare-and-exchange - that
    the compiler does /not/ know about, and is not supported by the C++ (or
    C) standard library for atomics.

    Use atomic_ref with a trivial structure that is 16 bytes with x64 or
    8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
    With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for inline-assembly.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Fri Sep 25 18:40:14 2026
    From Newsgroup: comp.lang.c++

    On 25/09/2026 18:25, Bonita Montero wrote:
    Am 21.09.2026 um 18:45 schrieb David Brown:

    The point here was an opcode - the 16 byte compare-and-exchange - that
    the compiler does /not/ know about, and is not supported by the C++
    (or C) standard library for atomics.

    Use atomic_ref with a trivial structure that is 16 bytes with x64 or
    8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
    With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for inline-assembly.

    My (admittedly very brief) testing suggested gcc calls a library
    function, rather than using the specific opcode directly in generated
    code. Perhaps this is affected by flags selecting the specific
    architecture. I neither know nor care particularly - that is beside the point. I am not interested in using these instructions, and leave that
    to the posters who want to implement RCU or other lock-free algorithms.

    Of course there are good reasons for using standard library methods for
    this kind of thing, rather than inline assembly, even if that results in slower code. But for those that have use of assembly, the post I wrote
    shows how it can be done for these example instructions, in a manner
    that is safer, more efficient and more flexible than the solutions
    provided by Joseph or Chris.

    I assume that they both know what is and is not supported by gcc and its current C++ standard library.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Joseph Seigh@jseigh_es00@xemaps.com to comp.lang.c++ on Fri Sep 25 13:47:59 2026
    From Newsgroup: comp.lang.c++

    On 9/25/26 12:25 PM, Bonita Montero wrote:
    Am 21.09.2026 um 18:45 schrieb David Brown:

    The point here was an opcode - the 16 byte compare-and-exchange - that
    the compiler does /not/ know about, and is not supported by the C++
    (or C) standard library for atomics.

    Use atomic_ref with a trivial structure that is 16 bytes with x64 or
    8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
    With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for inline-assembly.

    You need some way to avoid the ABA problem.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Fri Sep 25 20:54:32 2026
    From Newsgroup: comp.lang.c++

    On 25/09/2026 19:47, Joseph Seigh wrote:
    On 9/25/26 12:25 PM, Bonita Montero wrote:
    Am 21.09.2026 um 18:45 schrieb David Brown:

    The point here was an opcode - the 16 byte compare-and-exchange -
    that the compiler does /not/ know about, and is not supported by the
    C++ (or C) standard library for atomics.

    Use atomic_ref with a trivial structure that is 16 bytes with x64 or
    8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
    With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for
    inline-assembly.

    You need some way to avoid the ABA problem.


    That's entirely possible - and I expect you have more experience in that
    than I do. I have no experience with this on x86 - I work in microcontrollers, which are often simpler for such things. I haven't considered anything other than implementing the instruction - issues
    like ABA are no different when you use optimal inline assembly, less
    optimal inline assembly, or external assembly.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Joseph Seigh@jseigh_es00@xemaps.com to comp.lang.c++ on Fri Sep 25 17:20:43 2026
    From Newsgroup: comp.lang.c++

    On 9/25/26 2:54 PM, David Brown wrote:
    On 25/09/2026 19:47, Joseph Seigh wrote:
    On 9/25/26 12:25 PM, Bonita Montero wrote:

    Use atomic_ref with a trivial structure that is 16 bytes with x64 or
    8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
    With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for
    inline-assembly.

    You need some way to avoid the ABA problem.


    That's entirely possible - and I expect you have more experience in that than I do.-a I have no experience with this on x86 - I work in microcontrollers, which are often simpler for such things.-a I haven't considered anything other than implementing the instruction - issues
    like ABA are no different when you use optimal inline assembly, less
    optimal inline assembly, or external assembly.


    The 16 byte version lets you combine a memory reference w/ a number as a
    fat pointer to let you distinguish distinguish between two instances
    of an object sharing the same memory address (due to the memory being reallocated).* Basically the ABA problem.

    The example I posted is from a lock-free ring buffer, a bounded queue.
    Most of the lock-free ring buffers you see out there aren't really
    lock-free and/or an actual queue unless you qualify with "for some
    definition of lock-free" and "for some definition of queue".

    For linked queues, you have the Michael-Scott lock-free queue where
    you can get away with a single wide CAS but you need something like
    hazard pointers or RCU to avoid the ABA problem and to avoid doing
    CAS on a memory location with undefined state. LL/SC doesn't have
    this problem however.

    Double wide CAS support would be nice to have, but I'm not holding
    my breath.

    * In the 70's a double wide CAS on an IBM mainframe was 64 bits.
    Using a 32 bit number and 32 bit address, it was estimated that
    a 32 bit number would take about 100 years to wrap on the
    current hardware then.




    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Sat Sep 26 16:23:37 2026
    From Newsgroup: comp.lang.c++

    On 25/09/2026 23:20, Joseph Seigh wrote:
    On 9/25/26 2:54 PM, David Brown wrote:
    On 25/09/2026 19:47, Joseph Seigh wrote:
    On 9/25/26 12:25 PM, Bonita Montero wrote:

    Use atomic_ref with a trivial structure that is 16 bytes with x64 or
    8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
    With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for
    inline-assembly.

    You need some way to avoid the ABA problem.


    That's entirely possible - and I expect you have more experience in
    that than I do.-a I have no experience with this on x86 - I work in
    microcontrollers, which are often simpler for such things.-a I haven't
    considered anything other than implementing the instruction - issues
    like ABA are no different when you use optimal inline assembly, less
    optimal inline assembly, or external assembly.


    The 16 byte version lets you combine a memory reference w/ a number as a
    fat pointer to let you distinguish distinguish between two instances
    of an object sharing the same memory address (due to the memory being reallocated).*-a Basically the ABA problem.


    That fits my understand, which is nice to see.

    The example I posted is from a lock-free ring buffer, a bounded queue.
    Most of the lock-free ring buffers you see out there aren't really
    lock-free and/or an actual queue unless you qualify with "for some
    definition of lock-free" and "for some definition of queue".

    For linked queues, you have the Michael-Scott lock-free queue where
    you can get away with a single wide CAS but you need something like
    hazard pointers or RCU to avoid the ABA problem and to avoid doing
    CAS on a memory location with undefined state.-a LL/SC doesn't have
    this problem however.

    Double wide CAS support would be nice to have, but I'm not holding
    my breath.


    Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?

    * In the 70's a double wide CAS on an IBM mainframe was 64 bits.
    Using a 32 bit number and 32 bit address, it was estimated that
    a 32 bit number would take about 100 years to wrap on the
    current hardware then.





    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Joseph Seigh@jseigh_es00@xemaps.com to comp.lang.c++ on Sat Sep 26 15:36:59 2026
    From Newsgroup: comp.lang.c++

    On 9/26/26 10:23 AM, David Brown wrote:
    On 25/09/2026 23:20, Joseph Seigh wrote:


    Double wide CAS support would be nice to have, but I'm not holding
    my breath.


    Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?


    I meant as part of c/c++. There really no legitimate technical
    reason it's not in the standard now, it's just they can't put
    it in the way they would like too, so they refuse to put it in
    at all. The way they would like to is an atomic<T> where T would
    be an 8 byte type. But they can't since there are no atomic 8 byte
    load and store instructions on most existing x86-64 processors
    and that's the only way they are willing to do it.

    I have a pre c++11 implementation of atomic reference counting
    (actually atomic, not like Rust's ARC which is merely thread-safe)
    and there's no reason to rewrite it post c++11. I'd still need
    assembly.

    Anyway, there's solutions to more interesting concurrency problems to
    consider.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Sun Sep 27 14:56:21 2026
    From Newsgroup: comp.lang.c++

    On 26/09/2026 21:36, Joseph Seigh wrote:
    On 9/26/26 10:23 AM, David Brown wrote:
    On 25/09/2026 23:20, Joseph Seigh wrote:


    Double wide CAS support would be nice to have, but I'm not holding
    my breath.


    Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?


    I meant as part of c/c++.-a There really no legitimate technical
    reason it's not in the standard now, it's just they can't put
    it in the way they would like too, so they refuse to put it in
    at all.-a The way they would like to is an atomic<T> where T would
    be an 8 byte type.-a But they can't since there are no atomic 8 byte
    load and store instructions on most existing x86-64 processors
    and that's the only way they are willing to do it.


    Okay.

    There is a strong reluctance to put things into the standards if they
    can't be realistically and practically implemented on most platforms.

    But C++ /does/ support atomic<T> types, where T is 16-bytes - or any
    size at all. And it supports compare_exchange() methods (strong and
    weak) on those types.

    As far as I can see, it's up to the implementation how those are
    handled. gcc on x86-64 does so with a library call "__atomic_compare_exchange_16". (Bigger sizes use a more general __atomic_compare_exchange call.) If the target supports a dedicated instruction here, and the instruction is more efficient than the library
    call (that's usually the case!), then this is just a compiler quality of implementation issue. When someone adds this to the compiler and/or gcc
    C++ standard library implementation, it should just work when
    appropriate target selection flags are used.


    I have a pre c++11 implementation of atomic reference counting
    (actually atomic, not like Rust's ARC which is merely thread-safe)
    and there's no reason to rewrite it post c++11.-a I'd still need
    assembly.

    Anyway, there's solutions to more interesting concurrency problems to consider.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Mon Sep 28 15:42:13 2026
    From Newsgroup: comp.lang.c++

    On 9/24/2026 11:33 PM, David Brown wrote:
    On 25/09/2026 00:28, Chris M. Thomasson wrote:
    On 9/21/2026 5:50 AM, David Brown wrote:
    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- >>>>>>> sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -a-anever really got into inline asm. Well. I said f it and used good >>>> ol GAS. Or MASM.


    You really should learn about inline assembly in gcc.-a It can be
    significantly more efficient than using an external call in a
    separately assembled file.-a And I am confident that you'd find it
    fun!-a (Of course, [...]

    Yeah. Well. Shit. Back then I did not want any compiler to mess with
    things, and my inline skills were rather lacking! I felt way better
    making a header the structs, and externally assembling things. Keep in
    mind this was before C/C++ 11. I even said woe to LTO.


    I was not trying to show pre-C++11 Chris how to write inline assembly
    for atomics.-a I was trying to show 2026 Chris how to write inline
    assembly /today/, because it is still useful for some things (like a 16- byte compare-and-swap).-a I was also trying to show Joseph how to do a better job than he had so far.-a (Again, it's on the inline assembly
    part, not the choice of instructions or atomic handling.)

    You didn't know the details of gcc inline assembly back then - fair
    enough, it's not easy, and has subtly challenges.-a Now you know a bit
    more going forward.-a And I really do think it is something that you
    would enjoy playing with.


    Fwiw, I felt way better with externally assembled ASM back then. The
    syntax for GCC inline was/is a bit "bitter" to me, MASM was a little
    better, alas. But, I understand your main point.

    I still want to put it into a separate file, GAS it into an .o file and
    link the little shit.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Mon Sep 28 15:44:43 2026
    From Newsgroup: comp.lang.c++

    On 9/25/2026 2:20 PM, Joseph Seigh wrote:
    On 9/25/26 2:54 PM, David Brown wrote:
    On 25/09/2026 19:47, Joseph Seigh wrote:
    On 9/25/26 12:25 PM, Bonita Montero wrote:

    Use atomic_ref with a trivial structure that is 16 bytes with x64 or
    8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
    With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for
    inline-assembly.

    You need some way to avoid the ABA problem.


    That's entirely possible - and I expect you have more experience in
    that than I do.-a I have no experience with this on x86 - I work in
    microcontrollers, which are often simpler for such things.-a I haven't
    considered anything other than implementing the instruction - issues
    like ABA are no different when you use optimal inline assembly, less
    optimal inline assembly, or external assembly.


    The 16 byte version lets you combine a memory reference w/ a number as a
    fat pointer to let you distinguish distinguish between two instances
    of an object sharing the same memory address (due to the memory being reallocated).*-a Basically the ABA problem.

    The example I posted is from a lock-free ring buffer, a bounded queue.
    Most of the lock-free ring buffers you see out there aren't really
    lock-free and/or an actual queue unless you qualify with "for some
    definition of lock-free" and "for some definition of queue".

    For linked queues, you have the Michael-Scott lock-free queue where
    you can get away with a single wide CAS but you need something like
    hazard pointers or RCU to avoid the ABA problem and to avoid doing
    CAS on a memory location with undefined state.-a LL/SC doesn't have
    this problem however.

    Double wide CAS support would be nice to have, but I'm not holding
    my breath.

    Last time I checked god damn C++ said it was not lockfree wrt two
    contiguous words! So, I see why you are still using asm.



    * In the 70's a double wide CAS on an IBM mainframe was 64 bits.
    Using a 32 bit number and 32 bit address, it was estimated that
    a 32 bit number would take about 100 years to wrap on the
    current hardware then.





    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Mon Sep 28 15:45:38 2026
    From Newsgroup: comp.lang.c++

    On 9/25/2026 9:25 AM, Bonita Montero wrote:
    Am 21.09.2026 um 18:45 schrieb David Brown:

    The point here was an opcode - the 16 byte compare-and-exchange - that
    the compiler does /not/ know about, and is not supported by the C++
    (or C) standard library for atomics.

    Use atomic_ref with a trivial structure that is 16 bytes with x64 or
    8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
    With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for inline-assembly.

    Did they finally make it use the DWCAS and say its always lock-free?
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Mon Sep 28 15:50:53 2026
    From Newsgroup: comp.lang.c++

    On 9/27/2026 5:56 AM, David Brown wrote:
    On 26/09/2026 21:36, Joseph Seigh wrote:
    On 9/26/26 10:23 AM, David Brown wrote:
    On 25/09/2026 23:20, Joseph Seigh wrote:


    Double wide CAS support would be nice to have, but I'm not holding
    my breath.


    Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?


    I meant as part of c/c++.-a There really no legitimate technical
    reason it's not in the standard now, it's just they can't put
    it in the way they would like too, so they refuse to put it in
    at all.-a The way they would like to is an atomic<T> where T would
    be an 8 byte type.-a But they can't since there are no atomic 8 byte
    load and store instructions on most existing x86-64 processors
    and that's the only way they are willing to do it.


    Okay.

    There is a strong reluctance to put things into the standards if they
    can't be realistically and practically implemented on most platforms.

    But C++ /does/ support atomic<T> types, where T is 16-bytes - or any
    size at all.-a And it supports compare_exchange() methods (strong and
    weak) on those types.

    As far as I can see, it's up to the implementation how those are
    handled.-a gcc on x86-64 does so with a library call "__atomic_compare_exchange_16".-a (Bigger sizes use a more general __atomic_compare_exchange call.)-a If the target supports a dedicated instruction here, and the instruction is more efficient than the library call (that's usually the case!), then this is just a compiler quality of implementation issue.-a When someone adds this to the compiler and/or gcc C++ standard library implementation, it should just work when
    appropriate target selection flags are used.

    It is a QOI, as far as I can tell. Checked it a while ago, and it says
    that the DWCAS is not lock free. That scares me. If its not lock free,
    then there is no reason to use it for these types of things. ;^o





    I have a pre c++11 implementation of atomic reference counting
    (actually atomic, not like Rust's ARC which is merely thread-safe)
    and there's no reason to rewrite it post c++11.-a I'd still need
    assembly.

    Anyway, there's solutions to more interesting concurrency problems to
    consider.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Tue Sep 29 08:31:08 2026
    From Newsgroup: comp.lang.c++

    On 29/09/2026 00:50, Chris M. Thomasson wrote:
    On 9/27/2026 5:56 AM, David Brown wrote:
    On 26/09/2026 21:36, Joseph Seigh wrote:
    On 9/26/26 10:23 AM, David Brown wrote:
    On 25/09/2026 23:20, Joseph Seigh wrote:


    Double wide CAS support would be nice to have, but I'm not holding
    my breath.


    Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?


    I meant as part of c/c++.-a There really no legitimate technical
    reason it's not in the standard now, it's just they can't put
    it in the way they would like too, so they refuse to put it in
    at all.-a The way they would like to is an atomic<T> where T would
    be an 8 byte type.-a But they can't since there are no atomic 8 byte
    load and store instructions on most existing x86-64 processors
    and that's the only way they are willing to do it.


    Okay.

    There is a strong reluctance to put things into the standards if they
    can't be realistically and practically implemented on most platforms.

    But C++ /does/ support atomic<T> types, where T is 16-bytes - or any
    size at all.-a And it supports compare_exchange() methods (strong and
    weak) on those types.

    As far as I can see, it's up to the implementation how those are
    handled.-a gcc on x86-64 does so with a library call
    "__atomic_compare_exchange_16".-a (Bigger sizes use a more general
    __atomic_compare_exchange call.)-a If the target supports a dedicated
    instruction here, and the instruction is more efficient than the
    library call (that's usually the case!), then this is just a compiler
    quality of implementation issue.-a When someone adds this to the
    compiler and/or gcc C++ standard library implementation, it should
    just work when appropriate target selection flags are used.

    It is a QOI, as far as I can tell. Checked it a while ago, and it says
    that the DWCAS is not lock free. That scares me. If its not lock free,
    then there is no reason to use it for these types of things. ;^o


    The answer then would be to file an issue with the gcc bugzilla (since
    it's their C++ standard library implementation you are looking at).
    Feel free to post my inline assembly there in the issue, and post the
    issue link here.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Tue Sep 29 08:40:31 2026
    From Newsgroup: comp.lang.c++

    On 29/09/2026 00:42, Chris M. Thomasson wrote:
    On 9/24/2026 11:33 PM, David Brown wrote:
    On 25/09/2026 00:28, Chris M. Thomasson wrote:
    On 9/21/2026 5:50 AM, David Brown wrote:
    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-
    concurrent- sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -a-anever really got into inline asm. Well. I said f it and used good >>>>> ol GAS. Or MASM.


    You really should learn about inline assembly in gcc.-a It can be
    significantly more efficient than using an external call in a
    separately assembled file.-a And I am confident that you'd find it
    fun!-a (Of course, [...]

    Yeah. Well. Shit. Back then I did not want any compiler to mess with
    things, and my inline skills were rather lacking! I felt way better
    making a header the structs, and externally assembling things. Keep
    in mind this was before C/C++ 11. I even said woe to LTO.


    I was not trying to show pre-C++11 Chris how to write inline assembly
    for atomics.-a I was trying to show 2026 Chris how to write inline
    assembly /today/, because it is still useful for some things (like a
    16- byte compare-and-swap).-a I was also trying to show Joseph how to
    do a better job than he had so far.-a (Again, it's on the inline
    assembly part, not the choice of instructions or atomic handling.)

    You didn't know the details of gcc inline assembly back then - fair
    enough, it's not easy, and has subtly challenges.-a Now you know a bit
    more going forward.-a And I really do think it is something that you
    would enjoy playing with.


    Fwiw, I felt way better with externally assembled ASM back then. The
    syntax for GCC inline was/is a bit "bitter" to me, MASM was a little
    better, alas. But, I understand your main point.

    I still want to put it into a separate file, GAS it into an .o file and
    link the little shit.

    External assembly can be better for some types of code, but it is
    definitely inappropriate here. With external assembly you have to stick rigidly to the ABI for external calls, but it is clearer for large
    blocks of assembly (or other occasionally useful advanced stuff).
    However, the only reason you need large blocks of assembly for wrapping
    a single instruction is because you are using external assembly and the
    ABI for function calls! Inline assembly not only avoids that overhead,
    but lets the compiler integrate it much better with the rest of the code.

    The whole point of using lock-free algorithms here, and the double compare-and-swap, is speed. If you use external assembly so that all
    your important data has to be pushed from registers to the stack, then
    you have a call, then you pull the data off the stack again - you've a
    dozen extra instructions at the call site, a dozen extra instructions at
    the cas16 site, doing nothing except spoiling your efficient code.

    gcc inline assembly is a /much/ better solution here. It's okay if you
    don't know how to write it - and it is certainly okay if you did not
    know how to do so a decade ago. But you should understand why it is
    better, and then decide if you are going to learn about it.

    (Of course for this particular case, if you can persuade the gcc C++
    library folk to make their standard cas16 lock-free, then using the
    standard function is best.)





    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris M. Thomasson@chris.m.thomasson.1@gmail.com to comp.lang.c++ on Tue Sep 29 16:30:31 2026
    From Newsgroup: comp.lang.c++

    On 9/28/2026 11:40 PM, David Brown wrote:
    On 29/09/2026 00:42, Chris M. Thomasson wrote:
    On 9/24/2026 11:33 PM, David Brown wrote:
    On 25/09/2026 00:28, Chris M. Thomasson wrote:
    On 9/21/2026 5:50 AM, David Brown wrote:
    On 21/09/2026 01:09, Chris M. Thomasson wrote:
    On 9/20/2026 7:58 AM, Joseph Seigh wrote:
    On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
    On 9/16/2026 3:25 PM, Joseph Seigh wrote:
    https://jseigh.wordpress.com/2026/09/07/consequential-
    concurrent- sequential-rcu/


    C++ is sort of in the same boat, which is why we still need to
    use inline assembly to implement half century old algorithms
    efficiently.

    -a-anever really got into inline asm. Well. I said f it and used
    good ol GAS. Or MASM.


    You really should learn about inline assembly in gcc.-a It can be
    significantly more efficient than using an external call in a
    separately assembled file.-a And I am confident that you'd find it
    fun!-a (Of course, [...]

    Yeah. Well. Shit. Back then I did not want any compiler to mess with
    things, and my inline skills were rather lacking! I felt way better
    making a header the structs, and externally assembling things. Keep
    in mind this was before C/C++ 11. I even said woe to LTO.


    I was not trying to show pre-C++11 Chris how to write inline assembly
    for atomics.-a I was trying to show 2026 Chris how to write inline
    assembly /today/, because it is still useful for some things (like a
    16- byte compare-and-swap).-a I was also trying to show Joseph how to
    do a better job than he had so far.-a (Again, it's on the inline
    assembly part, not the choice of instructions or atomic handling.)

    You didn't know the details of gcc inline assembly back then - fair
    enough, it's not easy, and has subtly challenges.-a Now you know a bit
    more going forward.-a And I really do think it is something that you
    would enjoy playing with.


    Fwiw, I felt way better with externally assembled ASM back then. The
    syntax for GCC inline was/is a bit "bitter" to me, MASM was a little
    better, alas. But, I understand your main point.

    I still want to put it into a separate file, GAS it into an .o file
    and link the little shit.

    External assembly can be better for some types of code, but it is
    definitely inappropriate here.-a With external assembly you have to stick rigidly to the ABI for external calls, but it is clearer for large
    blocks of assembly (or other occasionally useful advanced stuff).
    However, the only reason you need large blocks of assembly for wrapping
    a single instruction is because you are using external assembly and the
    ABI for function calls!-a Inline assembly not only avoids that overhead,
    but lets the compiler integrate it much better with the rest of the code.

    The whole point of using lock-free algorithms here, and the double compare-and-swap, is speed.-a If you use external assembly so that all
    your important data has to be pushed from registers to the stack, then
    you have a call, then you pull the data off the stack again - you've a
    dozen extra instructions at the call site, a dozen extra instructions at
    the cas16 site, doing nothing except spoiling your efficient code.

    gcc inline assembly is a /much/ better solution here.-a It's okay if you don't know how to write it - and it is certainly okay if you did not
    know how to do so a decade ago.-a But you should understand why it is better, and then decide if you are going to learn about it.

    (Of course for this particular case, if you can persuade the gcc C++
    library folk to make their standard cas16 lock-free, then using the
    standard function is best.)
    Difficult to totally disagree with that David. Shit. Its my own personal problem with the syntax of GCC inline asm. I am a bit scared of it.
    Sigh. Clobbers what? When and where? cdecl for the ABI and externally assembled functions are fine for me, LTO aside for a moment, can be bad
    do not go into my asm! BUT. Again. Shit. Its a personal problem wrt me
    saying inline bad! Sorry!

    Also, if C++ gains a LOCK CMPXCHG16B for a double word, two contiguous
    64 bit words on a 64 bit machine, well GOOD! Show me. I have had some experiences with it trying to tell me its not lock free.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From David Brown@david.brown@hesbynett.no to comp.lang.c++ on Wed Sep 30 09:54:35 2026
    From Newsgroup: comp.lang.c++

    On 30/09/2026 01:30, Chris M. Thomasson wrote:
    On 9/28/2026 11:40 PM, David Brown wrote:

    gcc inline assembly is a /much/ better solution here.-a It's okay if
    you don't know how to write it - and it is certainly okay if you did
    not know how to do so a decade ago.-a But you should understand why it
    is better, and then decide if you are going to learn about it.

    (Of course for this particular case, if you can persuade the gcc C++
    library folk to make their standard cas16 lock-free, then using the
    standard function is best.)
    Difficult to totally disagree with that David. Shit. Its my own personal problem with the syntax of GCC inline asm. I am a bit scared of it.
    Sigh. Clobbers what? When and where? cdecl for the ABI and externally assembled functions are fine for me, LTO aside for a moment, can be bad
    do not go into my asm! BUT. Again. Shit. Its a personal problem wrt me saying inline bad! Sorry!


    Not knowing the syntax for gcc inline assembly is not a "problem" as
    such - most C programmers never need it. It is just a strong
    recommendation that /if/ you want to write assembly code, then your
    results would be much improved by using inline assembly format rather
    than external assembly functions. And you are already familiar with the
    hard part, the actual assembly. So don't look at this as a problem of
    any sort - look at it as a fun, challenging and useful new direction for
    your existing skills and interests.

    Also, if C++ gains a LOCK CMPXCHG16B for a double word, two contiguous
    64 bit words on a 64 bit machine, well GOOD! Show me. I have had some experiences with it trying to tell me its not lock free.

    The C++ standard library supports 16 byte compare-and-exchange. The gcc implementation of that library does not use that instruction, and thus
    does not support lock-free 16 byte compare-and-exchange. C++ has
    supported this since C++11 - this is a QOI issue, not a language or
    standard library issue. File an issue on the gcc bugzilla - that's how
    you get progress (or how we will learn if there are good reasons not to
    use it.)

    --- Synchronet 3.22a-Linux NewsLink 1.2