* tangentially related to something that I was working on which if far
more perverse but mostly in a solution that is looking for a problem
phase so it probably won't go anywhere.
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- sequential-rcu/
The name doesn't really mean anything. I was just generalizing something that's been around forever.*-a It's simple enough that it must be in use somewhere.-a I just haven't noticed it.-a I looked up the linux kernal seqlock documentation and it specifically says not to use pointers in
seqlock protected data and no idea why that is unless the linux kernel
is doing something weird with memory mapping.
Like seqlock, it is obstruction-free.-a If you have qsbr (RCU) or ebr,
that would be a better choice.-a It might be good for libraries where
the memory reclamation is hidden from the application.
* tangentially related to something that I was working on which if far
more perverse but mostly in a solution that is looking for a problem
phase so it probably won't go anywhere.
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
sequential-rcu/
The name doesn't really mean anything. I was just generalizing something
that's been around forever.*-a It's simple enough that it must be in use
somewhere.-a I just haven't noticed it.-a I looked up the linux kernal
seqlock documentation and it specifically says not to use pointers in
seqlock protected data and no idea why that is unless the linux kernel
is doing something weird with memory mapping.
Like seqlock, it is obstruction-free.-a If you have qsbr (RCU) or ebr,
that would be a better choice.-a It might be good for libraries where
the memory reclamation is hidden from the application.
* tangentially related to something that I was working on which if far
more perverse but mostly in a solution that is looking for a problem
phase so it probably won't go anywhere.
Fwiw, I have used seqlocks for the writer side and pure RCU for the
readers. Actually, there is a way to use seqlocks to gain DWCAS on a
system that does not support it.
https://groups.google.com/g/lock-free/c/X3fuuXknQF0/m/zfHnoFi-VXgJ
https://pastebin.com/raw/TgTcfYtR
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
sequential-rcu/
The name doesn't really mean anything. I was just generalizing something >>> that's been around forever.*-a It's simple enough that it must be in use >>> somewhere.-a I just haven't noticed it.-a I looked up the linux kernal
seqlock documentation and it specifically says not to use pointers in
seqlock protected data and no idea why that is unless the linux kernel
is doing something weird with memory mapping.
Like seqlock, it is obstruction-free.-a If you have qsbr (RCU) or ebr,
that would be a better choice.-a It might be good for libraries where
the memory reclamation is hidden from the application.
* tangentially related to something that I was working on which if far
more perverse but mostly in a solution that is looking for a problem
phase so it probably won't go anywhere.
Fwiw, I have used seqlocks for the writer side and pure RCU for the
readers. Actually, there is a way to use seqlocks to gain DWCAS on a
system that does not support it.
https://groups.google.com/g/lock-free/c/X3fuuXknQF0/m/zfHnoFi-VXgJ
https://pastebin.com/raw/TgTcfYtR
Java did its AtomicStampedReference by creating an object to hold
the stamp and the reference and then doing a single word compare
and swap on a reference to it.-a All this so you can run on
hardware which effectively doesn't exist anymore and even if it
did you would need to backport the os to support obsolete
hardware so the jvm could even run on it.
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
Re the OP, I did think of a really good POC for this but it is
way too much work.-a I have better things to waste my time
with. :)
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
sequential-rcu/
The name doesn't really mean anything. I was just generalizing something >>> that's been around forever.*-a It's simple enough that it must be in use >>> somewhere.-a I just haven't noticed it.-a I looked up the linux kernal
seqlock documentation and it specifically says not to use pointers in
seqlock protected data and no idea why that is unless the linux kernel
is doing something weird with memory mapping.
Like seqlock, it is obstruction-free.-a If you have qsbr (RCU) or ebr,
that would be a better choice.-a It might be good for libraries where
the memory reclamation is hidden from the application.
* tangentially related to something that I was working on which if far
more perverse but mostly in a solution that is looking for a problem
phase so it probably won't go anywhere.
Fwiw, I have used seqlocks for the writer side and pure RCU for the
readers. Actually, there is a way to use seqlocks to gain DWCAS on a
system that does not support it.
https://groups.google.com/g/lock-free/c/X3fuuXknQF0/m/zfHnoFi-VXgJ
https://pastebin.com/raw/TgTcfYtR
Java did its AtomicStampedReference by creating an object to hold
the stamp and the reference and then doing a single word compare
and swap on a reference to it.-a All this so you can run on
hardware which effectively doesn't exist anymore and even if it
did you would need to backport the os to support obsolete
hardware so the jvm could even run on it.
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
Re the OP, I did think of a really good POC for this but it is
way too much work.-a I have better things to waste my time
with. :)
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-anever really got into inline asm. Well. I said f it and used good ol
GAS. Or MASM.
GAS.
https://web.archive.org/web/20060214112345/http:// appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_gcc_asm.html
https://web.archive.org/web/20060214112539/http:// appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_masm_asm.html
GAS was kind to me.
On 9/20/26 7:09 PM, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-a-anever really got into inline asm. Well. I said f it and used good ol
GAS. Or MASM.
GAS.
https://web.archive.org/web/20060214112345/http://
appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_gcc_asm.html
https://web.archive.org/web/20060214112539/http://
appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_masm_asm.html
GAS was kind to me.
https://github.com/jseigh/queues/blob/main/include/atomix.h
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-anever really got into inline asm. Well. I said f it and used good ol
GAS. Or MASM.
https://web.archive.org/web/20060214112345/http:// appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_gcc_asm.html
https://web.archive.org/web/20060214112539/http:// appcore.home.comcast.net/appcore/src/cpu/i686/ac_i686_masm_asm.html
GAS was kind to me.
# Copyright 2005 Chris Thomasson
.align 16
.globl np_ac_i686_atomic_dwcas_fence
np_ac_i686_atomic_dwcas_fence:
-a pushl %esi
-a pushl %ebx
-a movl 16(%esp), %esi
-a movl (%esi), %eax
-a movl 4(%esi), %edx
-a movl 20(%esp), %esi
-a movl (%esi), %ebx
-a movl 4(%esi), %ecx
-a movl 12(%esp), %esi
-a lock cmpxchg8b (%esi)
-a jne np_ac_i686_atomic_dwcas_fence_fail
-a xorl %eax, %eax
-a popl %ebx
-a popl %esi
-a ret
np_ac_i686_atomic_dwcas_fence_fail:
-a movl 16(%esp), %esi
-a movl %eax, (%esi)
-a movl %edx, 4(%esi)
-a movl $1, %eax
-a popl %ebx
-a popl %esi
ret
On 21/09/2026 01:09, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-anever really got into inline asm. Well. I said f it and used good ol
GAS. Or MASM.
You really should learn about inline assembly in gcc. It can be >significantly more efficient than using an external call in a separately >assembled file. And I am confident that you'd find it fun! (Of course,
On Mon, 21 Sep 2026 14:50:19 +0200
David Brown <david.brown@hesbynett.no> gabbled:
On 21/09/2026 01:09, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- >>>>>> sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-a-anever really got into inline asm. Well. I said f it and used good
ol GAS. Or MASM.
You really should learn about inline assembly in gcc.-a It can be
significantly more efficient than using an external call in a
separately assembled file.-a And I am confident that you'd find it
fun!-a (Of course,
Unless the assembler is just a few opcodes you're better off leaving it to
a compiler as it'll usually do a better job and probably knows opcodes you've
never even heard of.
On 21/09/2026 01:09, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent-
sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-a-anever really got into inline asm. Well. I said f it and used good ol
GAS. Or MASM.
You really should learn about inline assembly in gcc.-a It can be significantly more efficient than using an external call in a separately assembled file.-a And I am confident that you'd find it fun!-a (Of course, [...]
On Mon, 21 Sep 2026 14:50:19 +0200
David Brown <david.brown@hesbynett.no> gabbled:
On 21/09/2026 01:09, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- >>>>>> sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-a-anever really got into inline asm. Well. I said f it and used good
ol GAS. Or MASM.
You really should learn about inline assembly in gcc.-a It can be
significantly more efficient than using an external call in a
separately assembled file.-a And I am confident that you'd find it
fun!-a (Of course,
Unless the assembler is just a few opcodes you're better off leaving it to
a compiler as it'll usually do a better job and probably knows opcodes you've
never even heard of.
On 9/21/2026 5:50 AM, David Brown wrote:
On 21/09/2026 01:09, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- >>>>>> sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-a-anever really got into inline asm. Well. I said f it and used good
ol GAS. Or MASM.
You really should learn about inline assembly in gcc.-a It can be
significantly more efficient than using an external call in a
separately assembled file.-a And I am confident that you'd find it
fun!-a (Of course, [...]
Yeah. Well. Shit. Back then I did not want any compiler to mess with
things, and my inline skills were rather lacking! I felt way better
making a header the structs, and externally assembling things. Keep in
mind this was before C/C++ 11. I even said woe to LTO.
The point here was an opcode - the 16 byte compare-and-exchange - that
the compiler does /not/ know about, and is not supported by the C++ (or
C) standard library for atomics.
Am 21.09.2026 um 18:45 schrieb David Brown:
The point here was an opcode - the 16 byte compare-and-exchange - that
the compiler does /not/ know about, and is not supported by the C++
(or C) standard library for atomics.
Use atomic_ref with a trivial structure that is 16 bytes with x64 or
8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for inline-assembly.
Am 21.09.2026 um 18:45 schrieb David Brown:
The point here was an opcode - the 16 byte compare-and-exchange - that
the compiler does /not/ know about, and is not supported by the C++
(or C) standard library for atomics.
Use atomic_ref with a trivial structure that is 16 bytes with x64 or
8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for inline-assembly.
On 9/25/26 12:25 PM, Bonita Montero wrote:
Am 21.09.2026 um 18:45 schrieb David Brown:
The point here was an opcode - the 16 byte compare-and-exchange -
that the compiler does /not/ know about, and is not supported by the
C++ (or C) standard library for atomics.
Use atomic_ref with a trivial structure that is 16 bytes with x64 or
8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for
inline-assembly.
You need some way to avoid the ABA problem.
On 25/09/2026 19:47, Joseph Seigh wrote:
On 9/25/26 12:25 PM, Bonita Montero wrote:
Use atomic_ref with a trivial structure that is 16 bytes with x64 or
8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for
inline-assembly.
You need some way to avoid the ABA problem.
That's entirely possible - and I expect you have more experience in that than I do.-a I have no experience with this on x86 - I work in microcontrollers, which are often simpler for such things.-a I haven't considered anything other than implementing the instruction - issues
like ABA are no different when you use optimal inline assembly, less
optimal inline assembly, or external assembly.
On 9/25/26 2:54 PM, David Brown wrote:
On 25/09/2026 19:47, Joseph Seigh wrote:
On 9/25/26 12:25 PM, Bonita Montero wrote:
The 16 byte version lets you combine a memory reference w/ a number as aUse atomic_ref with a trivial structure that is 16 bytes with x64 or
8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for
inline-assembly.
You need some way to avoid the ABA problem.
That's entirely possible - and I expect you have more experience in
that than I do.-a I have no experience with this on x86 - I work in
microcontrollers, which are often simpler for such things.-a I haven't
considered anything other than implementing the instruction - issues
like ABA are no different when you use optimal inline assembly, less
optimal inline assembly, or external assembly.
fat pointer to let you distinguish distinguish between two instances
of an object sharing the same memory address (due to the memory being reallocated).*-a Basically the ABA problem.
The example I posted is from a lock-free ring buffer, a bounded queue.
Most of the lock-free ring buffers you see out there aren't really
lock-free and/or an actual queue unless you qualify with "for some
definition of lock-free" and "for some definition of queue".
For linked queues, you have the Michael-Scott lock-free queue where
you can get away with a single wide CAS but you need something like
hazard pointers or RCU to avoid the ABA problem and to avoid doing
CAS on a memory location with undefined state.-a LL/SC doesn't have
this problem however.
Double wide CAS support would be nice to have, but I'm not holding
my breath.
* In the 70's a double wide CAS on an IBM mainframe was 64 bits.
Using a 32 bit number and 32 bit address, it was estimated that
a 32 bit number would take about 100 years to wrap on the
current hardware then.
On 25/09/2026 23:20, Joseph Seigh wrote:
Double wide CAS support would be nice to have, but I'm not holding
my breath.
Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?
On 9/26/26 10:23 AM, David Brown wrote:
On 25/09/2026 23:20, Joseph Seigh wrote:
Double wide CAS support would be nice to have, but I'm not holding
my breath.
Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?
I meant as part of c/c++.-a There really no legitimate technical
reason it's not in the standard now, it's just they can't put
it in the way they would like too, so they refuse to put it in
at all.-a The way they would like to is an atomic<T> where T would
be an 8 byte type.-a But they can't since there are no atomic 8 byte
load and store instructions on most existing x86-64 processors
and that's the only way they are willing to do it.
I have a pre c++11 implementation of atomic reference counting
(actually atomic, not like Rust's ARC which is merely thread-safe)
and there's no reason to rewrite it post c++11.-a I'd still need
assembly.
Anyway, there's solutions to more interesting concurrency problems to consider.
On 25/09/2026 00:28, Chris M. Thomasson wrote:
On 9/21/2026 5:50 AM, David Brown wrote:
On 21/09/2026 01:09, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-concurrent- >>>>>>> sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-a-anever really got into inline asm. Well. I said f it and used good >>>> ol GAS. Or MASM.
You really should learn about inline assembly in gcc.-a It can be
significantly more efficient than using an external call in a
separately assembled file.-a And I am confident that you'd find it
fun!-a (Of course, [...]
Yeah. Well. Shit. Back then I did not want any compiler to mess with
things, and my inline skills were rather lacking! I felt way better
making a header the structs, and externally assembling things. Keep in
mind this was before C/C++ 11. I even said woe to LTO.
I was not trying to show pre-C++11 Chris how to write inline assembly
for atomics.-a I was trying to show 2026 Chris how to write inline
assembly /today/, because it is still useful for some things (like a 16- byte compare-and-swap).-a I was also trying to show Joseph how to do a better job than he had so far.-a (Again, it's on the inline assembly
part, not the choice of instructions or atomic handling.)
You didn't know the details of gcc inline assembly back then - fair
enough, it's not easy, and has subtly challenges.-a Now you know a bit
more going forward.-a And I really do think it is something that you
would enjoy playing with.
On 9/25/26 2:54 PM, David Brown wrote:
On 25/09/2026 19:47, Joseph Seigh wrote:
On 9/25/26 12:25 PM, Bonita Montero wrote:
The 16 byte version lets you combine a memory reference w/ a number as aUse atomic_ref with a trivial structure that is 16 bytes with x64 or
8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for
inline-assembly.
You need some way to avoid the ABA problem.
That's entirely possible - and I expect you have more experience in
that than I do.-a I have no experience with this on x86 - I work in
microcontrollers, which are often simpler for such things.-a I haven't
considered anything other than implementing the instruction - issues
like ABA are no different when you use optimal inline assembly, less
optimal inline assembly, or external assembly.
fat pointer to let you distinguish distinguish between two instances
of an object sharing the same memory address (due to the memory being reallocated).*-a Basically the ABA problem.
The example I posted is from a lock-free ring buffer, a bounded queue.
Most of the lock-free ring buffers you see out there aren't really
lock-free and/or an actual queue unless you qualify with "for some
definition of lock-free" and "for some definition of queue".
For linked queues, you have the Michael-Scott lock-free queue where
you can get away with a single wide CAS but you need something like
hazard pointers or RCU to avoid the ABA problem and to avoid doing
CAS on a memory location with undefined state.-a LL/SC doesn't have
this problem however.
Double wide CAS support would be nice to have, but I'm not holding
my breath.
* In the 70's a double wide CAS on an IBM mainframe was 64 bits.
Using a 32 bit number and 32 bit address, it was estimated that
a 32 bit number would take about 100 years to wrap on the
current hardware then.
Am 21.09.2026 um 18:45 schrieb David Brown:
The point here was an opcode - the 16 byte compare-and-exchange - that
the compiler does /not/ know about, and is not supported by the C++
(or C) standard library for atomics.
Use atomic_ref with a trivial structure that is 16 bytes with x64 or
8 bytes with x86. This works with MSVC, g++ and clang++ on x86 / x64.
With that you have CMPXCHG16B (x64) or CMPXCHG8B (x86). No need for inline-assembly.
On 26/09/2026 21:36, Joseph Seigh wrote:
On 9/26/26 10:23 AM, David Brown wrote:
On 25/09/2026 23:20, Joseph Seigh wrote:
Double wide CAS support would be nice to have, but I'm not holding
my breath.
Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?
I meant as part of c/c++.-a There really no legitimate technical
reason it's not in the standard now, it's just they can't put
it in the way they would like too, so they refuse to put it in
at all.-a The way they would like to is an atomic<T> where T would
be an 8 byte type.-a But they can't since there are no atomic 8 byte
load and store instructions on most existing x86-64 processors
and that's the only way they are willing to do it.
Okay.
There is a strong reluctance to put things into the standards if they
can't be realistically and practically implemented on most platforms.
But C++ /does/ support atomic<T> types, where T is 16-bytes - or any
size at all.-a And it supports compare_exchange() methods (strong and
weak) on those types.
As far as I can see, it's up to the implementation how those are
handled.-a gcc on x86-64 does so with a library call "__atomic_compare_exchange_16".-a (Bigger sizes use a more general __atomic_compare_exchange call.)-a If the target supports a dedicated instruction here, and the instruction is more efficient than the library call (that's usually the case!), then this is just a compiler quality of implementation issue.-a When someone adds this to the compiler and/or gcc C++ standard library implementation, it should just work when
appropriate target selection flags are used.
I have a pre c++11 implementation of atomic reference counting
(actually atomic, not like Rust's ARC which is merely thread-safe)
and there's no reason to rewrite it post c++11.-a I'd still need
assembly.
Anyway, there's solutions to more interesting concurrency problems to
consider.
On 9/27/2026 5:56 AM, David Brown wrote:
On 26/09/2026 21:36, Joseph Seigh wrote:
On 9/26/26 10:23 AM, David Brown wrote:
On 25/09/2026 23:20, Joseph Seigh wrote:
Double wide CAS support would be nice to have, but I'm not holding
my breath.
Doesn't the cmpxchg16b count as a double-wide CAS on x86-64 ?
I meant as part of c/c++.-a There really no legitimate technical
reason it's not in the standard now, it's just they can't put
it in the way they would like too, so they refuse to put it in
at all.-a The way they would like to is an atomic<T> where T would
be an 8 byte type.-a But they can't since there are no atomic 8 byte
load and store instructions on most existing x86-64 processors
and that's the only way they are willing to do it.
Okay.
There is a strong reluctance to put things into the standards if they
can't be realistically and practically implemented on most platforms.
But C++ /does/ support atomic<T> types, where T is 16-bytes - or any
size at all.-a And it supports compare_exchange() methods (strong and
weak) on those types.
As far as I can see, it's up to the implementation how those are
handled.-a gcc on x86-64 does so with a library call
"__atomic_compare_exchange_16".-a (Bigger sizes use a more general
__atomic_compare_exchange call.)-a If the target supports a dedicated
instruction here, and the instruction is more efficient than the
library call (that's usually the case!), then this is just a compiler
quality of implementation issue.-a When someone adds this to the
compiler and/or gcc C++ standard library implementation, it should
just work when appropriate target selection flags are used.
It is a QOI, as far as I can tell. Checked it a while ago, and it says
that the DWCAS is not lock free. That scares me. If its not lock free,
then there is no reason to use it for these types of things. ;^o
On 9/24/2026 11:33 PM, David Brown wrote:
On 25/09/2026 00:28, Chris M. Thomasson wrote:
On 9/21/2026 5:50 AM, David Brown wrote:
On 21/09/2026 01:09, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-
concurrent- sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-a-anever really got into inline asm. Well. I said f it and used good >>>>> ol GAS. Or MASM.
You really should learn about inline assembly in gcc.-a It can be
significantly more efficient than using an external call in a
separately assembled file.-a And I am confident that you'd find it
fun!-a (Of course, [...]
Yeah. Well. Shit. Back then I did not want any compiler to mess with
things, and my inline skills were rather lacking! I felt way better
making a header the structs, and externally assembling things. Keep
in mind this was before C/C++ 11. I even said woe to LTO.
I was not trying to show pre-C++11 Chris how to write inline assembly
for atomics.-a I was trying to show 2026 Chris how to write inline
assembly /today/, because it is still useful for some things (like a
16- byte compare-and-swap).-a I was also trying to show Joseph how to
do a better job than he had so far.-a (Again, it's on the inline
assembly part, not the choice of instructions or atomic handling.)
You didn't know the details of gcc inline assembly back then - fair
enough, it's not easy, and has subtly challenges.-a Now you know a bit
more going forward.-a And I really do think it is something that you
would enjoy playing with.
Fwiw, I felt way better with externally assembled ASM back then. The
syntax for GCC inline was/is a bit "bitter" to me, MASM was a little
better, alas. But, I understand your main point.
I still want to put it into a separate file, GAS it into an .o file and
link the little shit.
On 29/09/2026 00:42, Chris M. Thomasson wrote:Difficult to totally disagree with that David. Shit. Its my own personal problem with the syntax of GCC inline asm. I am a bit scared of it.
On 9/24/2026 11:33 PM, David Brown wrote:
On 25/09/2026 00:28, Chris M. Thomasson wrote:
On 9/21/2026 5:50 AM, David Brown wrote:
On 21/09/2026 01:09, Chris M. Thomasson wrote:
On 9/20/2026 7:58 AM, Joseph Seigh wrote:
On 9/19/26 4:34 PM, Chris M. Thomasson wrote:
On 9/16/2026 3:25 PM, Joseph Seigh wrote:
https://jseigh.wordpress.com/2026/09/07/consequential-
concurrent- sequential-rcu/
C++ is sort of in the same boat, which is why we still need to
use inline assembly to implement half century old algorithms
efficiently.
-a-anever really got into inline asm. Well. I said f it and used
good ol GAS. Or MASM.
You really should learn about inline assembly in gcc.-a It can be
significantly more efficient than using an external call in a
separately assembled file.-a And I am confident that you'd find it
fun!-a (Of course, [...]
Yeah. Well. Shit. Back then I did not want any compiler to mess with
things, and my inline skills were rather lacking! I felt way better
making a header the structs, and externally assembling things. Keep
in mind this was before C/C++ 11. I even said woe to LTO.
I was not trying to show pre-C++11 Chris how to write inline assembly
for atomics.-a I was trying to show 2026 Chris how to write inline
assembly /today/, because it is still useful for some things (like a
16- byte compare-and-swap).-a I was also trying to show Joseph how to
do a better job than he had so far.-a (Again, it's on the inline
assembly part, not the choice of instructions or atomic handling.)
You didn't know the details of gcc inline assembly back then - fair
enough, it's not easy, and has subtly challenges.-a Now you know a bit
more going forward.-a And I really do think it is something that you
would enjoy playing with.
Fwiw, I felt way better with externally assembled ASM back then. The
syntax for GCC inline was/is a bit "bitter" to me, MASM was a little
better, alas. But, I understand your main point.
I still want to put it into a separate file, GAS it into an .o file
and link the little shit.
External assembly can be better for some types of code, but it is
definitely inappropriate here.-a With external assembly you have to stick rigidly to the ABI for external calls, but it is clearer for large
blocks of assembly (or other occasionally useful advanced stuff).
However, the only reason you need large blocks of assembly for wrapping
a single instruction is because you are using external assembly and the
ABI for function calls!-a Inline assembly not only avoids that overhead,
but lets the compiler integrate it much better with the rest of the code.
The whole point of using lock-free algorithms here, and the double compare-and-swap, is speed.-a If you use external assembly so that all
your important data has to be pushed from registers to the stack, then
you have a call, then you pull the data off the stack again - you've a
dozen extra instructions at the call site, a dozen extra instructions at
the cas16 site, doing nothing except spoiling your efficient code.
gcc inline assembly is a /much/ better solution here.-a It's okay if you don't know how to write it - and it is certainly okay if you did not
know how to do so a decade ago.-a But you should understand why it is better, and then decide if you are going to learn about it.
(Of course for this particular case, if you can persuade the gcc C++
library folk to make their standard cas16 lock-free, then using the
standard function is best.)
On 9/28/2026 11:40 PM, David Brown wrote:
gcc inline assembly is a /much/ better solution here.-a It's okay ifDifficult to totally disagree with that David. Shit. Its my own personal problem with the syntax of GCC inline asm. I am a bit scared of it.
you don't know how to write it - and it is certainly okay if you did
not know how to do so a decade ago.-a But you should understand why it
is better, and then decide if you are going to learn about it.
(Of course for this particular case, if you can persuade the gcc C++
library folk to make their standard cas16 lock-free, then using the
standard function is best.)
Sigh. Clobbers what? When and where? cdecl for the ABI and externally assembled functions are fine for me, LTO aside for a moment, can be bad
do not go into my asm! BUT. Again. Shit. Its a personal problem wrt me saying inline bad! Sorry!
Also, if C++ gains a LOCK CMPXCHG16B for a double word, two contiguous
64 bit words on a 64 bit machine, well GOOD! Show me. I have had some experiences with it trying to tell me its not lock free.
| Sysop: | Amessyroom |
|---|---|
| Location: | Fayetteville, NC |
| Users: | 74 |
| Nodes: | 6 (0 / 6) |
| Uptime: | 121:05:53 |
| Calls: | 1,194 |
| Files: | 1,352 |
| Messages: | 290,204 |