From Newsgroup: sci.physics
Hi,
He uses FIFO, and DMA and Noc:
Getting peak TOPS on a Ryzen AI 7 350 NPU
https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/
But lets say whether its FIFO or FILO
isn't so importand his used cases are,
what is now found in my library(furryhaze)
for GPU, namely the very basic:
/**
* test_gpu_comp_start(W, K): internal only
* The predicate succeeds. As a side effect it
* starts the -C-WAM W with K warps.
*/
function test_gpu_comp_start(args)
/**
* test_gpu_comp_join(W, P): internal only
* The predicate succeeds in P with a new promise
* that waits for the -C-WAM W to finish.
*/
function test_gpu_comp_join(args)
A GPU interface, via the command processor
for example of WebGPU, does the above
synchronization for you.
In the NPU example he does everything
low level, with Python IRON an stuff:
"Since the main way to achieve synchronization
within the IRON framework is by doing data
movement with object FIFOs, IrCOm sending a
dummy uint32 value as some sort of
synchronization token.
Waiting for all the kernels to finish is
trickier. The object FIFOs support a join
pattern in which an object FIFO consumes an
object from each of multiple object FIFOs,
concatenates these objects and produces the
concatenated object as a result.
Etc.."
Getting peak TOPS on a Ryzen AI 7 350 NPU
https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/
So Daniel Est|-vez Scientific & Technical
Amateur Radio, gives a nice glimpse into an
NPU, I have not yet publicitly released
my library(furryhaze), since its still in
testing. Maybe take another week or so,
still I have ironed out all corners,
for example the new gpu_comp_start and
gpu_comp_join works fine on may desktop
AI laptops, but I have still a bug on
my iPad AI tablet, on the Redmi AI phone,
also chokes on a test case.
Bye
Mild Shock schrieb:
Hi,
You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:
From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Chris M. Thomasson schrieb:
I don't think he knows exactly what he is doing...
Why does he need a lock/wait-free queue in a
compute shader? What is he trying to do?
--- Synchronet 3.22a-Linux NewsLink 1.2