Hi,
Because of MIMD you have to reassess algorithms.
A spin loop which could really hurt non-MIMD
GPUs, might less hurt a MIMD GPU.
Basically you have to reassess algorithms. Be
very exact whether your claims relates to
non-MIMD or to MIMD. You can toy around
with WebGL, mainly made for non-MIMD, here:
https://www.shadertoy.com/
and with WegGPU, mainly made for MIMD, here:
https://compute.toys/
The compute toys page, supports two shader
languages, WGSL and Slang. I made my GPU
experiments only with WGSL.
So I don't know Slang either. Things I
don't know in the GPU world are:
- OpenGL 4.2 and later
- Slang https://shader-slang.org/
Things I have meanwhile hands on, and which
I plan to integrate into library(edge/brainfog):
- WGSL https://webgpufundamentals.org/
Bye
Mild Shock schrieb:
Hi,
I went to holdays in June 2026, had an idea for
a pi-WAM based on a Hack, the later is described here:
Emulating -C-WAM in Dogelog Player
https://medium.com/2989/de9cd29c7d37
The Elements of Computing Systems
https://mitpress.mit.edu/9780262539807
In July 2026 I did the CPU and GPU experiments,
moving from emulator to native executor based in
realizing Hack as a concrete virtual machine,
and not as an abstract machine emulated in Prolog.
The GPU experiments were done in WebGPU / WGSL.
So no, I never used OpenGL Version 4.2 and later.
Also the name imageAtomicAdd indicates that
imageAtomicAdd is rather from a render shader,
while my GPU experiment uses a compute shader.
Especially I need GPU compute shaders, which
are not executed in lock step, but rather have
indepdendent thread state, also known as MIMD.
"In computing, multiple instruction, multiple
data (MIMD) is a technique employed to
achieve parallelism. "
https://en.wikipedia.org/wiki/Multiple_instruction,_multiple_data
MIMID showed up 2017 with NVIDIA Volta cards.
But is now realized by Intel Arc, Snapdragon Adreno,
AMD RDNA and Apple Silicon as well.
Bye
Chris M. Thomasson schrieb:
On 7/20/2026 4:32 PM, Mild Shock wrote:
Hi,[...]
imageAtomicAdd is trivial, but it does
not help with bounded buffers.
Are you sure about that! ;^o
How about Silicon Grid Engine and MPI, OpenMP and old cluster.
Or old "batch jobs".
Batch jobs:-a it's how work gets done.
You crazy frothing lunatic
Hi,
Rossy Boy is slower than Micro Penis.
Both being heavy alcoholic. Both don't
understand the "budget" here:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Budget means , I don't use a Mainframe
with a Job Control language. Budget means
ca. 1000 USD for the AI Laptops (Yoga, Ryzen
and Think) I bought end of 2025, and ca. 500
USD for the AI Laptop (Mac Neo) I bought
middle of 2026. I explained that already to
the Micro Penis moron. Now I explain it
again to the Rossy Boy herpes blister
corona victim.
Bye
Ross Finlayson schrieb:
How about Silicon Grid Engine and MPI, OpenMP and old cluster.
Or old "batch jobs".
Batch jobs:-a it's how work gets done.
You crazy frothing lunatic
Hi,
But please go on Rossy Boy, ask more stupid
questions, that I have already answered in
relation to Micro Penis.
I am happy to repeat, what I already have
posted, again and again. If necesssary I
will repeat the well known material,
that everybody can find on the internet,
again like 1000x times. I have no problem with
that providing this information again and
again, as long as stupid questions are asked.
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is slower than Micro Penis.
Both being heavy alcoholic. Both don't
understand the "budget" here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Budget means , I don't use a Mainframe
with a Job Control language. Budget means
ca. 1000 USD for the AI Laptops (Yoga, Ryzen
and Think) I bought end of 2025, and ca. 500
USD for the AI Laptop (Mac Neo) I bought
middle of 2026. I explained that already to
the Micro Penis moron. Now I explain it
again to the Rossy Boy herpes blister
corona victim.
Bye
Ross Finlayson schrieb:
How about Silicon Grid Engine and MPI, OpenMP and old cluster.
Or old "batch jobs".
Batch jobs:-a it's how work gets done.
You crazy frothing lunatic
Hi,
I don't have the feeling its rocket science
what I did. But maybe, since idiocracy has
definitively reached computer science,
you can now call yourself software engineer,
with a Hackathon certificate, and a GitHub
account. I might be mistaken, and you need
to be Einstain to understand the Giga Lips result?
Idiocracy (2006) - Movie Trailer
https://www.youtube.com/watch?v=te5vtOEz7sY
"Two things are infinite: the universe and human
stupidity; and I'm not sure about the universe."
-- Einstein
Bye
Mild Shock schrieb:
Hi,
But please go on Rossy Boy, ask more stupid
questions, that I have already answered in
relation to Micro Penis.
I am happy to repeat, what I already have
posted, again and again. If necesssary I
will repeat the well known material,
that everybody can find on the internet,
again like 1000x times. I have no problem with
that providing this information again and
again, as long as stupid questions are asked.
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is slower than Micro Penis.
Both being heavy alcoholic. Both don't
understand the "budget" here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Budget means , I don't use a Mainframe
with a Job Control language. Budget means
ca. 1000 USD for the AI Laptops (Yoga, Ryzen
and Think) I bought end of 2025, and ca. 500
USD for the AI Laptop (Mac Neo) I bought
middle of 2026. I explained that already to
the Micro Penis moron. Now I explain it
again to the Rossy Boy herpes blister
corona victim.
Bye
Ross Finlayson schrieb:
How about Silicon Grid Engine and MPI, OpenMP and old cluster.
Or old "batch jobs".
Batch jobs:-a it's how work gets done.
You crazy frothing lunatic
Frameworks like WebLLM leverage WebGPU to run large language models
locally inside browsers like Google Chrome, providing completely
private, offline AI assistants.
Hi,
But please go on Rossy Boy, ask more stupid
questions, that I have already answered in
relation to Micro Penis.
I am happy to repeat, what I already have
posted, again and again. If necesssary I
will repeat the well known material,
that everybody can find on the internet,
again like 1000x times. I have no problem with
that providing this information again and
again, as long as stupid questions are asked.
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is slower than Micro Penis.
Both being heavy alcoholic. Both don't
understand the "budget" here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Budget means , I don't use a Mainframe
with a Job Control language. Budget means
ca. 1000 USD for the AI Laptops (Yoga, Ryzen
and Think) I bought end of 2025, and ca. 500
USD for the AI Laptop (Mac Neo) I bought
middle of 2026. I explained that already to
the Micro Penis moron. Now I explain it
again to the Rossy Boy herpes blister
corona victim.
Bye
Ross Finlayson schrieb:
How about Silicon Grid Engine and MPI, OpenMP and old cluster.
Or old "batch jobs".
Batch jobs: it's how work gets done.
You crazy frothing lunatic
You crazy frothing lunatic
Mild Shock wrote:
Frameworks like WebLLM leverage WebGPU to run large language models
locally inside browsers like Google Chrome, providing completely
private, offline AI assistants.
since when google chrome private, think again
Hi,
Nice Freundian Slip the below self
description of yours. Rossy Boy!
LoL
Bye
Ross Finlayson schrieb:
You crazy frothing lunatic
Hi,
------------------- begin --------------------
Teaching Micro Penis Vilage Idiot
------------------- begin --------------------
You dont have to use WebLLM, respectively
WebGPU / WGSL literally, just read the next
post I did AND use your brains moron:
Like WebAssembly before it, WebGPU has "escaped" the browser https://blog.4dpipeline.com/client-side-ai-is-here-how-webgpu-transforms-your-gpu-server-economics
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically access
an offline AI assistant, via some REST end-point
on your machine. Nothing to do with Google Chrome
browser security. You can make it as private as you want, by
having a firewall and not outward or inward
connection at all, only your REST end-point
on your machine. Or if you want a REST end-point
on a server of yours in the same intranet.
You don't need to use the internet, or put
something on the extranet, or use some sort of
subscription. What you need is access through
the firewall to download the REST software
and the LLM model. Tools like LM Studio and oMLX
offer this download and also install REST endpoint.
------------------- end --------------------
Teaching Micro Penis Vilage Idiot}
------------------- end --------------------
Bye
Will Bakshandaev schrieb:
Mild Shock wrote:
Frameworks like WebLLM leverage WebGPU to run large language models
locally inside browsers like Google Chrome, providing completely
private, offline AI assistants.
since when google chrome private, think again
Windows Recall takes a screenshot of a user'shttps://en.wikipedia.org/wiki/Windows_Recall
desktop every few seconds, then uses on-device
large language models to allow a user to
retrieve items and information that had
previously been on their screen.
Hi,
Currently companies such as Apple, Windows, etc..
are hardning their operating systems, so
that they can provide agentic AI sandboxes.
Problem is an agentic AI, that acts on your
behalf, when not enough supervised, might
do all kind of stuff on its own. So how do you
have harder borders. Besides companies that
write operating systems, there is also a cottage
industry now that adresses this paranoia,
here an example from a former Prologer:
Stop guessing what your coding agent just did
Prempti: Guardrails and Observability for AI Coding Agents. https://prempti.falco.org/
IntelliJ doesn't have this problem, it shows a
not yet hyper locally commited change, in the editor,
created by the AI, that you can review, and
then hyper locally commit in the editor. Only then
it lands in the file system. But also there it
will be subject to the local history and repository
version system. So the IntelliJ AI is pretty smartly
implemented, and hooks into their editors and newly
introduced hyper change visualization, a feature that
probably codemirror doesn't have yet. Have to double check.
Have Fun!
Bye
Mild Shock schrieb:
Hi,
------------------- begin --------------------
Teaching Micro Penis Vilage Idiot
------------------- begin --------------------
You dont have to use WebLLM, respectively
WebGPU / WGSL literally, just read the next
post I did AND use your brains moron:
Like WebAssembly before it, WebGPU has "escaped" the browser
https://blog.4dpipeline.com/client-side-ai-is-here-how-webgpu-transforms-your-gpu-server-economics
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically access
an offline AI assistant, via some REST end-point
on your machine. Nothing to do with Google Chrome
browser security. You can make it as private as you want, by
having a firewall and not outward or inward
connection at all, only your REST end-point
on your machine. Or if you want a REST end-point
on a server of yours in the same intranet.
You don't need to use the internet, or put
something on the extranet, or use some sort of
subscription. What you need is access through
the firewall to download the REST software
and the LLM model. Tools like LM Studio and oMLX
offer this download and also install REST endpoint.
------------------- end --------------------
Teaching Micro Penis Vilage Idiot}
------------------- end --------------------
Bye
Will Bakshandaev schrieb:
Mild Shock wrote:
Frameworks like WebLLM leverage WebGPU to run large language models
locally inside browsers like Google Chrome, providing completely
private, offline AI assistants.
since when google chrome private, think again
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically access an offline AI assistant, via some REST end-point
on your machine. Nothing to do with Google Chrome browser security. You
can make it as private as you want, by having a firewall and not outward
or inward
Mild Shock wrote:
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically
access an offline AI assistant, via some REST end-point
yet one more proof this half german inbreed
is an imbecile, ports go up to 16bits/64k only,
idiot, you cant have a localhost: whatever wrong
number you put there. You extreme fucking idiot.
on your machine. Nothing to do with Google
Chrome browser security. You can make it as private
as you want, by having a firewall and not outward
or inward
yes, i can see your point, they just want
your private cellphone number,
there rest is private and free, idiot
Hi,
------------------- begin --------------------
Teaching Micro Penis Vilage Idiot
------------------- begin --------------------
You dont have to use WebLLM, respectively
WebGPU / WGSL literally, just read the next
post I did AND use your brains moron:
Like WebAssembly before it, WebGPU has "escaped" the browser https://blog.4dpipeline.com/client-side-ai-is-here-how-webgpu-transforms-your-gpu-server-economics
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically access
an offline AI assistant, via some REST end-point
on your machine. Nothing to do with Google Chrome
browser security. You can make it as private as you want, by
having a firewall and not outward or inward
connection at all, only your REST end-point
on your machine. Or if you want a REST end-point
on a server of yours in the same intranet.
You don't need to use the internet, or put
something on the extranet, or use some sort of
subscription. What you need is access through
the firewall to download the REST software
and the LLM model. Tools like LM Studio and oMLX
offer this download and also install REST endpoint.
------------------- end --------------------
Teaching Micro Penis Vilage Idiot}
------------------- end --------------------
Bye
Will Bakshandaev schrieb:
Mild Shock wrote:
Frameworks like WebLLM leverage WebGPU to run large language models
locally inside browsers like Google Chrome, providing completely
private, offline AI assistants.
since when google chrome private, think again
Hi,
Maybe you could post some subtantial critique moron?
Instead of gibberish all the time. What does a cellphone
number have to do with a REST endpoint? Nothing!
Its all locally and my laptop has no cellphone number:
Start the REST API server
To start the server, run the following command:
lms server start
Endpoints
GET /api/v0/models
List all loaded and downloaded models
Example request
curl -H "Authorization: Bearer $LM_API_TOKEN" http://localhost:1234/api/v0/models
Response format
{
-a "object": "list",
-a "data": [
-a-a-a {
-a-a-a-a-a "id": "qwen2-vl-7b-instruct",
-a-a-a-a-a "object": "model",
-a-a-a-a-a "type": "vlm",
-a-a-a-a-a "publisher": "mlx-community",
-a-a-a-a-a "arch": "qwen2_vl"
Etc...
https://lmstudio.ai/docs/developer/rest/endpoints
Bye
BTW: LM Studio recently introduced LM Link,
which provides some VPN. It can be used to
create clients or servers that run models.
It is end-to-end encrypted, and built on top
of custom Tailscale mesh VPNs. This is for
the paranoid, that want to acccess a
LLM from one end of the globe, that sits
on the other end of the globe, and have
no evesdroper or whatever on the
information that is exchanged.
Mild Shock wrote:
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically access an offline AI
assistant, via some REST end-point
yet one more proof this half german inbreed is an imbecile, ports go
up to 16bits/64k only, idiot, you cant have a localhost: whatever
wrong number you put there. You extreme fucking idiot.
on your machine. Nothing to do with Google Chrome browser security.
You can make it as private as you want, by having a firewall and not
outward
or inward
yes, i can see your point, they just want your private cellphone
number, there rest is private and free, idiot
Mild Shock schrieb:
Hi,
------------------- begin --------------------
Teaching Micro Penis Vilage Idiot
------------------- begin --------------------
You dont have to use WebLLM, respectively
WebGPU / WGSL literally, just read the next
post I did AND use your brains moron:
Like WebAssembly before it, WebGPU has "escaped" the browser
https://blog.4dpipeline.com/client-side-ai-is-here-how-webgpu-transforms-your-gpu-server-economics
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically access
an offline AI assistant, via some REST end-point
on your machine. Nothing to do with Google Chrome
browser security. You can make it as private as you want, by
having a firewall and not outward or inward
connection at all, only your REST end-point
on your machine. Or if you want a REST end-point
on a server of yours in the same intranet.
You don't need to use the internet, or put
something on the extranet, or use some sort of
subscription. What you need is access through
the firewall to download the REST software
and the LLM model. Tools like LM Studio and oMLX
offer this download and also install REST endpoint.
------------------- end --------------------
Teaching Micro Penis Vilage Idiot}
------------------- end --------------------
Bye
Will Bakshandaev schrieb:
Mild Shock wrote:
Frameworks like WebLLM leverage WebGPU to run large language models
locally inside browsers like Google Chrome, providing completely
private, offline AI assistants.
since when google chrome private, think again
Hi,
How confused is tiny winy penis?
For the 100th time the budget here:
11.4 Giga Lips with a Budget Laptop
Ryzen AI 7 350 w/ Radeon 860M https://github.com/Jean-Luc-Picard-2021/gigabudget
is a laptop and not a smartphone. It
has no cellphone number. And w/ means
integrated GPU on the silicon chip,
and not a GPU connected to the mainboard
via some PCI bus. The model is a acer
swift go, I already posted this info:
Swift Go 16 AI SFG16-61-R21J Notebook https://www.acer.com/ch-de/laptops/swift/swift-go-16-ai-amd/pdp/NX.JCREZ.007
Bye
Mild Shock schrieb:
Hi,
Maybe you could post some subtantial critique moron?
Instead of gibberish all the time. What does a cellphone
number have to do with a REST endpoint? Nothing!
Its all locally and my laptop has no cellphone number:
Start the REST API server
To start the server, run the following command:
lms server start
Endpoints
GET /api/v0/models
List all loaded and downloaded models
Example request
curl -H "Authorization: Bearer $LM_API_TOKEN"
http://localhost:1234/api/v0/models
Response format
{
-a-a "object": "list",
-a-a "data": [
-a-a-a-a {
-a-a-a-a-a-a "id": "qwen2-vl-7b-instruct",
-a-a-a-a-a-a "object": "model",
-a-a-a-a-a-a "type": "vlm",
-a-a-a-a-a-a "publisher": "mlx-community",
-a-a-a-a-a-a "arch": "qwen2_vl"
Etc...
https://lmstudio.ai/docs/developer/rest/endpoints
Bye
BTW: LM Studio recently introduced LM Link,
which provides some VPN. It can be used to
create clients or servers that run models.
It is end-to-end encrypted, and built on top
of custom Tailscale mesh VPNs. This is for
the paranoid, that want to acccess a
LLM from one end of the globe, that sits
on the other end of the globe, and have
no evesdroper or whatever on the
information that is exchanged.
Mild Shock wrote:
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically access an offline AI
assistant, via some REST end-point
yet one more proof this half german inbreed is an imbecile, ports go
up to 16bits/64k only, idiot, you cant have a localhost: whatever
wrong number you put there. You extreme fucking idiot.
on your machine. Nothing to do with Google Chrome browser security.
You can make it as private as you want, by having a firewall and not
outward
or inward
yes, i can see your point, they just want your private cellphone
number, there rest is private and free, idiot
Mild Shock schrieb:
Hi,
------------------- begin --------------------
Teaching Micro Penis Vilage Idiot
------------------- begin --------------------
You dont have to use WebLLM, respectively
WebGPU / WGSL literally, just read the next
post I did AND use your brains moron:
Like WebAssembly before it, WebGPU has "escaped" the browser
https://blog.4dpipeline.com/client-side-ai-is-here-how-webgpu-transforms-your-gpu-server-economics
http://localhost:567921/ is private you moron.,
or what ever port REST is using. You typically access
an offline AI assistant, via some REST end-point
on your machine. Nothing to do with Google Chrome
browser security. You can make it as private as you want, by
having a firewall and not outward or inward
connection at all, only your REST end-point
on your machine. Or if you want a REST end-point
on a server of yours in the same intranet.
You don't need to use the internet, or put
something on the extranet, or use some sort of
subscription. What you need is access through
the firewall to download the REST software
and the LLM model. Tools like LM Studio and oMLX
offer this download and also install REST endpoint.
------------------- end --------------------
Teaching Micro Penis Vilage Idiot}
------------------- end --------------------
Bye
Will Bakshandaev schrieb:
Mild Shock wrote:
Frameworks like WebLLM leverage WebGPU to run large language models
locally inside browsers like Google Chrome, providing completely
private, offline AI assistants.
since when google chrome private, think again
Maybe you could post some subtantial critique moron? Instead of
gibberish all the time. What does a cellphone number have to do with a
REST endpoint? Nothing!
Its all locally and my laptop has no cellphone number:
you stupid half german, the comparisons along
memories arrays gpu card located, are taking place
parallel without cpu intervention, you fucking
illiterate idiot. You soon will become a quarter
german hence 3/4 russian old days, historically. It's coming
German energy crisis caused by rCylack of Russian gasrCO rCo Merz https://www.rt.com/news/643254-germany-crisis-russian-gas/
Mild Shock wrote:
Maybe you could post some subtantial critique moron? Instead of
gibberish all the time. What does a cellphone number have to do with a
REST endpoint? Nothing!
Its all locally and my laptop has no cellphone number:
you lying bitch, your arse is burning, you said google chrome security,
and gave them your cellphone number, fucking idiot. And you put wrong
ports numbers along the localhost: idiot
Moron, from St. Petersburg, with only 5G internet.
You confuse browser security model with Google account.
No cellphone number involved in a brower JavaScript
How confused is tiny winy penis?
For the 100th time the budget here:
11.4 Giga Lips with a Budget Laptop
Ryzen AI 7 350 w/ Radeon 860M https://github.com/Jean-Luc-Picard-2021/gigabudget
is a laptop and not a smartphone. It
has no cellphone number. And w/ means
integrated GPU on the silicon chip,
and not a GPU connected to the mainboard
via some PCI bus. The model is a acer
swift go, I already posted this info:
Swift Go 16 AI SFG16-61-R21J Notebook https://www.acer.com/ch-de/laptops/swift/swift-go-16-ai-amd/pdp/NX.JCREZ.007
Mild Shock wrote:
Moron, from St. Petersburg, with only 5G internet.
You confuse browser security model with Google account.
No cellphone number involved in a brower JavaScript
cretin, they already have your phone number, the IMEI, email adr, location and everything, your friends included. Idiot, I cant even believe it.
Hi,
Die IMEI (International Mobile Equipment Identity)
ist eine 15-stellige, weltweit eindeutige Seriennummer,
die jedes Mobiltelefon identifiziert.
Die meisten Laptops haben keine IMEI-Nummer. Sie
existiert nur, wenn das Ger|nt |+ber ein integriertes
Mobilfunkmodem (WWAN/LTE/5G-Karte) f|+r SIM-Karten verf|+gt.
Also in Switzerland people have fiber optic earth
cables for internet, and not some 5G over the air.
I got like 10 GBit/s fiber here in my home.
There are 96 internet service providers that offer that speed:
https://www.comparis.ch/telecom/zuhause/angebote/list?requestobject={%22products%22%3A[1]%2C%22onlyOffersWithoutMinimumDuration%22%3Afalse%2C%22internetSpeedTypes%22%3A[1]%2C%22connectionTypes%22%3A[1]%2C%22tvOptions%22%3A[]%2C%22landlineOptions%22%3A[]%2C%22providers%22%3A[]%2C%22showDiscountedOnly%22%3Afalse%2C%22addressCheckInfo%22%3A{%22AvailableCableProviders%22%3Anull%2C%22AvailableVdslSpeed%22%3Anull%2C%22AvailableFiberSpeed%22%3Anull%2C%22AvailableInit7Fiber%22%3Afalse%2C%22AvailableAllFiberProvidersExceptInit7%22%3Afalse}%2C%22address%22%3A{%22Zip%22%3Anull%2C%22Street%22%3Anull%2C%22StreetNumber%22%3Anull}}
Angebote f|+r Internet, TV und Festnetz-Telefon, sowie Kombiangebote
Its handy to download LLMs which have GB sizes.
But this laptop has no SIM Card, even not a e-SIM:
How confused is tiny winy penis?
For the 100th time the budget here:
11.4 Giga Lips with a Budget Laptop
Ryzen AI 7 350 w/ Radeon 860M
https://github.com/Jean-Luc-Picard-2021/gigabudget
is a laptop and not a smartphone. It
has no cellphone number. And w/ means
integrated GPU on the silicon chip,
and not a GPU connected to the mainboard
via some PCI bus. The model is a acer
swift go, I already posted this info:
Swift Go 16 AI SFG16-61-R21J Notebook
https://www.acer.com/ch-de/laptops/swift/swift-go-16-ai-amd/pdp/NX.JCREZ.007
Bye
Roque Bahtinov schrieb:
Mild Shock wrote:
Moron, from St. Petersburg, with only 5G internet.
You confuse browser security model with Google account.
No cellphone number involved in a brower JavaScript
cretin, they already have your phone number, the IMEI, email adr,
location
and everything, your friends included. Idiot, I cant even believe it.
Also in Switzerland people have fiber optic earth cables for internet,
and not some 5G over the air.
I got like 10 GBit/s fiber here in my home.
Mild Shock wrote:
Also in Switzerland people have fiber optic earth cables for internet,
and not some 5G over the air.
I got like 10 GBit/s fiber here in my home.
wow, you think they are stupid, they cant see your location through the fiber; you are like a search tree, idiot, they know the name and the ID of all the crap you have around; then you said google chrome is private,
private my ass
Since I don't belong to some post CCCP, cigaret smuggling cartell, I
have nothing to hide. If you are clever you find
I have nothing to hide, you can find me in search.ch
Mild Shock wrote:
Since I don't belong to some post CCCP, cigaret smuggling cartell, I
have nothing to hide. If you are clever you find
hence you are admitting you are a fucking inbreed half german idiot from birth. Thanks making it clearer.
Hi,
I have nothing to hide, you can find me in search.ch
Since when belongs .ch to German? The .ch
is the country code top-level domain (ccTLD)
for Switzerland in the Domain Name System
of the Internet.
If you want find proof of Germany, you
would need to see .de. That I use a German
usenet provider doesn't mean I am German.
Everybody can use solani.org.
Switzerland, SWITCH
https://en.wikipedia.org/wiki/.ch
Germany, DENIC
https://en.wikipedia.org/wiki/.de
nicht-kommerziellen Usenet-News-Server
https://solani.org/
Whats wrong with you?
Bye
Keiv Babenchikov schrieb:
Mild Shock wrote:
Since I don't belong to some post CCCP, cigaret smuggling cartell, I
have nothing to hide. If you are clever you find
hence you are admitting you are a fucking inbreed half german idiot from
birth. Thanks making it clearer.
Since the Vodka has burn all your brain cells, ask a Ukrainian Neighbour
to do Detective.
You might find people smarter than you, you are already at the lower
left end in the Gauss
curve, of an IQ suitable for software engineering.
On 07/20/2026 11:59 PM, Mild Shock wrote:
Hi,
Because of MIMD you have to reassess algorithms.
A spin loop which could really hurt non-MIMD
GPUs, might less hurt a MIMD GPU.
Basically you have to reassess algorithms. Be
very exact whether your claims relates to
non-MIMD or to MIMD. You can toy around
with WebGL, mainly made for non-MIMD, here:
https://www.shadertoy.com/
and with WegGPU, mainly made for MIMD, here:
https://compute.toys/
The compute toys page, supports two shader
languages, WGSL and Slang. I made my GPU
experiments only with WGSL.
So I don't know Slang either. Things I
don't know in the GPU world are:
- OpenGL 4.2 and later
- Slang https://shader-slang.org/
Things I have meanwhile hands on, and which
I plan to integrate into library(edge/brainfog):
- WGSL https://webgpufundamentals.org/
Bye
Mild Shock schrieb:
Hi,
I went to holdays in June 2026, had an idea for
a pi-WAM based on a Hack, the later is described here:
Emulating -C-WAM in Dogelog Player
https://medium.com/2989/de9cd29c7d37
The Elements of Computing Systems
https://mitpress.mit.edu/9780262539807
In July 2026 I did the CPU and GPU experiments,
moving from emulator to native executor based in
realizing Hack as a concrete virtual machine,
and not as an abstract machine emulated in Prolog.
The GPU experiments were done in WebGPU / WGSL.
So no, I never used OpenGL Version 4.2 and later.
Also the name imageAtomicAdd indicates that
imageAtomicAdd is rather from a render shader,
while my GPU experiment uses a compute shader.
Especially I need GPU compute shaders, which
are not executed in lock step, but rather have
indepdendent thread state, also known as MIMD.
"In computing, multiple instruction, multiple
data (MIMD) is a technique employed to
achieve parallelism. "
https://en.wikipedia.org/wiki/Multiple_instruction,_multiple_data
MIMID showed up 2017 with NVIDIA Volta cards.
But is now realized by Intel Arc, Snapdragon Adreno,
AMD RDNA and Apple Silicon as well.
Bye
Chris M. Thomasson schrieb:
On 7/20/2026 4:32 PM, Mild Shock wrote:
Hi,[...]
imageAtomicAdd is trivial, but it does
not help with bounded buffers.
Are you sure about that! ;^o
How about Silicon Grid Engine and MPI, OpenMP and old cluster.
Or old "batch jobs".
Batch jobs:-a it's how work gets done.
You crazy frothing lunatic
Hi,
CAS and XADD have no looping, they
are atomic operations, that take some
time but basically have some outcome
Mild Shock wrote:
Hi,
CAS and XADD have no looping, they
are atomic operations, that take some
time but basically have some outcome
My father says C++ reminds him of an abortion.
Hi,
Your father is regreting not using a
contraceptive. Now there is just one more
moron walking earth, and that moron
is you Lane W alias micro penis.
Bye
Lane W schrieb:
Mild Shock wrote:
Hi,
CAS and XADD have no looping, they
are atomic operations, that take some
time but basically have some outcome
My father says C++ reminds him of an abortion.
Mild Shock wrote:
Hi,Look at C++. It's awful. C# is ten times better. There isn't even type string in C++. I could never go back to that.
Your father is regreting not using a
contraceptive. Now there is just one more
moron walking earth, and that moron
is you Lane W alias micro penis.
Bye
Lane W schrieb:
Mild Shock wrote:
Hi,
CAS and XADD have no looping, they
are atomic operations, that take some
time but basically have some outcome
My father says C++ reminds him of an abortion.
while (!enqueue(q, val)) ; /** Looping **/
}
}
Do you see the two loops, in your C code and in my Java code? They are
marked with a comment /** Looping **/ .
Look at C++. It's awful. C# is ten times better. There isn't even typeMy father says C++ reminds him of an abortion.
string in C++. I could never go back to that.
Mild Shock wrote:
while (!enqueue(q, val)) ; /** Looping **/
}
}
Do you see the two loops, in your C code and in my Java code? They are
marked with a comment /** Looping **/ .
that's a while(FALSE) idiot, this guy doesnt know what he has there
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
I don't think he knows exactly what he is doing... Why does he need a lock/wait-free queue in a compute shader? What is he trying to do?
Hi,
You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:
From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Chris M. Thomasson schrieb:
I don't think he knows exactly what he is doing... Why does he need a
lock/wait-free queue in a compute shader? What is he trying to do?
Hi,
You can also deduce that I need comms,
from pi in pi-WAM, since pi refers to pi-calculus.
There is also a nice paper, that I have already posted:
A pi-calculus Specification of Prolog https://scispace.com/pdf/a-pi-calculus-specification-of-prolog-3qf2pf04ud.pdf
You asked yourself why no atomic and
only comms? So its as simple as 1+1=2.
But usenet people are usually slow as fuck.
Take your time. You could spin loop to ingest
the topic, i.e. try again in 3-4 months, for
example reading some of the paper. Although
I know thats a totally unrealistic request, asking
a troll to do RTFM and study something. They
rather make themselves a total laughing stock,
play stupid games, win usenet prizes.
Bye
Mild Shock schrieb:
Hi,
You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:
From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Chris M. Thomasson schrieb:
I don't think he knows exactly what he is doing... Why does he need a
lock/wait-free queue in a compute shader? What is he trying to do?
Hi,
You can also deduce that I need comms,
from pi in pi-WAM, since pi refers to pi-calculus.
There is also a nice paper, that I have already posted:
Show an outline of what you need you compute shader to do?
On 7/24/2026 5:52 AM, Mild Shock wrote:
Hi,
You can also deduce that I need comms,
from pi in pi-WAM, since pi refers to pi-calculus.
There is also a nice paper, that I have already posted:
Show an outline of what you need you compute shader to do? I know about them. My code loves to saturate points during iteration of some fun
things I am working on. BUT! I need those points to accumulate. So, I
use fetch-and-add in the compute shader to make sure that the
accumulations are coherent.
Its 100% loopless. Strive for avoid loops at all costs if you can,
epically in the GPU.
[...]
Hi,
You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:
From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Chris M. Thomasson schrieb:
I don't think he knows exactly what he is doing... Why does he need a
lock/wait-free queue in a compute shader? What is he trying to do?
On 07/24/2026 05:43 AM, Mild Shock wrote:
Hi,
You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:
From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Chris M. Thomasson schrieb:
I don't think he knows exactly what he is doing... Why does he need a
lock/wait-free queue in a compute shader? What is he trying to do?
Forget Space-X and forget that heil-throwing fat-ass, too,
the rockets they bought are wearing out and the ones they
built are blowing up, moving fast and breaking things.
Shut Up
Hi,
The primary motivation behind SpaceX's acquisition
of Cursor (and its underlying Composer model tech stack)
comes down to vertical integration of the entire AI
stackrCocompute, models, and application surface.
Bridging the Coding & Reasoning Gap: While xAIrCOs Grok
had massive raw compute backing it (via massive clusters
like Colossus), its training data heavily leaned on
social media feeds from X.
That made it great for conversational chat, but left
it lagging behind competitors (like OpenAI and Anthropic)
in pure coding, reasoning efficiency, and long-horizon
agentic execution.
Owning the Developer Workflow: Instead of just building
a model and hoping people use it, acquiring Cursor
gives SpaceX the premier developer-facing application
layer that millions of developers already use
for daily coding.
Coupling Cursor's model architecture with SpaceX/xAI's
massive compute infrastructure allows them to aggressively
scale up training for next-gen models.
So its a Win-Win. Got it?
Bye
Ross Finlayson schrieb:
On 07/24/2026 05:43 AM, Mild Shock wrote:
Hi,
You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:
From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Chris M. Thomasson schrieb:
I don't think he knows exactly what he is doing... Why does he need a
lock/wait-free queue in a compute shader? What is he trying to do?
Forget Space-X and forget that heil-throwing fat-ass, too,
the rockets they bought are wearing out and the ones they
built are blowing up, moving fast and breaking things.
Shut Up
On 07/24/2026 04:23 PM, Mild Shock wrote:
Hi,
The primary motivation behind SpaceX's acquisition
of Cursor (and its underlying Composer model tech stack)
comes down to vertical integration of the entire AI
stackrCocompute, models, and application surface.
Bridging the Coding & Reasoning Gap: While xAIrCOs Grok
had massive raw compute backing it (via massive clusters
like Colossus), its training data heavily leaned on
social media feeds from X.
That made it great for conversational chat, but left
it lagging behind competitors (like OpenAI and Anthropic)
in pure coding, reasoning efficiency, and long-horizon
agentic execution.
Owning the Developer Workflow: Instead of just building
a model and hoping people use it, acquiring Cursor
gives SpaceX the premier developer-facing application
layer that millions of developers already use
for daily coding.
Coupling Cursor's model architecture with SpaceX/xAI's
massive compute infrastructure allows them to aggressively
scale up training for next-gen models.
So its a Win-Win. Got it?
Bye
Ross Finlayson schrieb:
On 07/24/2026 05:43 AM, Mild Shock wrote:
Hi,
You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:
From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Chris M. Thomasson schrieb:
I don't think he knows exactly what he is doing... Why does he need a >>>>> lock/wait-free queue in a compute shader? What is he trying to do?
Forget Space-X and forget that heil-throwing fat-ass, too,
the rockets they bought are wearing out and the ones they
built are blowing up, moving fast and breaking things.
Shut Up
I detest when people say "Got it" I think it's ridiculous,
like "unnerstand". It's immature and insensate.
If there was a functioning free-market capitalism,
there would be modular, safe, fun, fast vehicles of
the electric, hybrid, and liquid fueled varieties
that were open, reliable, and geared toward
owner/operators, not tyro turds.
Douchebag
Forget Space-X and forget that heil-throwing fat-ass, too,
the rockets they bought are wearing out and the ones they
built are blowing up, moving fast and breaking things.
Hi,
You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:
From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Chris M. Thomasson schrieb:
I don't think he knows exactly what he is doing... Why does he need a
lock/wait-free queue in a compute shader? What is he trying to do?
Then not to waste GPU compute time, by logical
threads doing nothing.
On 7/24/2026 5:43 AM, Mild Shock wrote:
Then not to waste GPU compute time, by logical
threads doing nothing.
Well, then never get to a full/empty condition. It depends on what you
are trying to do. If a GPU thread, warp needs to wait on something, then
you are not designing things right to begin with?
On 8/1/2026 2:02 AM, Chris M. Thomasson wrote:
On 7/24/2026 5:43 AM, Mild Shock wrote:
Then not to waste GPU compute time, by logical
threads doing nothing.
Well, then never get to a full/empty condition. It depends on what you
are trying to do. If a GPU thread, warp needs to wait on something,
then you are not designing things right to begin with?
Are you sure you even need FIFO? There is a really fast LIFO stack that
is also atomic. Now, for the GPU you should never have to spinwait, or
wait on anything. You need to be able to always have work to do. There
are certian patterns that work well. Its not like on the CPU where we
can wait in the kernel on conditions, ala futex or something.
Hi,
He uses FIFO, and DMA and Noc:
Getting peak TOPS on a Ryzen AI 7 350 NPU https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/
But lets say whether its FIFO or FILO
isn't so importand his used cases are,
what is now found in my library(furryhaze)
for GPU, namely the very basic:
/**
-a* test_gpu_comp_start(W, K): internal only
-a* The predicate succeeds. As a side effect it
-a* starts the -C-WAM W with K warps.
-a*/
function test_gpu_comp_start(args)
/**
-a* test_gpu_comp_join(W, P): internal only
-a* The predicate succeeds in P with a new promise
-a* that waits for the -C-WAM W to finish.
-a*/
function test_gpu_comp_join(args)
A GPU interface, via the command processor
for example of WebGPU, does the above
synchronization for you.
In the NPU example he does everything
low level, with Python IRON an stuff:
"Since the main way to achieve synchronization
within the IRON framework is by doing data
movement with object FIFOs, IrCOm sending a
dummy uint32 value as some sort of
synchronization token.
Waiting for all the kernels to finish is
trickier. The object FIFOs support a join
pattern in which an object FIFO consumes an
object from each of multiple object FIFOs,
concatenates these objects and produces the
concatenated object as a result.
Etc.."
Getting peak TOPS on a Ryzen AI 7 350 NPU https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/
So Daniel Est|-vez Scientific & Technical
Amateur Radio, gives a nice glimpse into an
NPU, I have not yet publicitly released
my library(furryhaze), since its still in
testing. Maybe take another week or so,
still I have ironed out all corners,
for example the new gpu_comp_start and
gpu_comp_join works fine on may desktop
AI laptops, but I have still a bug on
my iPad AI tablet, on the Redmi AI phone,
also chokes on a test case.
Bye
Chris M. Thomasson schrieb:
On 8/1/2026 2:02 AM, Chris M. Thomasson wrote:
On 7/24/2026 5:43 AM, Mild Shock wrote:
Then not to waste GPU compute time, by logical
threads doing nothing.
Well, then never get to a full/empty condition. It depends on what
you are trying to do. If a GPU thread, warp needs to wait on
something, then you are not designing things right to begin with?
Are you sure you even need FIFO? There is a really fast LIFO stack
that is also atomic. Now, for the GPU you should never have to
spinwait, or wait on anything. You need to be able to always have work
to do. There are certian patterns that work well. Its not like on the
CPU where we can wait in the kernel on conditions, ala futex or
something.
Hi,[...]
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
Of course we can make a special texture to handle it.
On 8/1/2026 5:22 AM, Mild Shock wrote:
Hi,[...]
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You are going to need a place to allocate nodes in the compute shader.
Of course we can make a special texture to handle it. But, we need to
strive to avoid a wait condition. I don't want a compute shader to spin. Yes, CAS can be used, but, try to make it be used as a "state machine", where the transitions from states are atomic. Try to avoid it making a
loop, where we loop on failure.
Hi,
WebGPU and WebGL are two different things. I explained
that towards you already like 3-5 times.
Of course we can make a special texture to handle it.
You still don't understand that I am using WebGPU,
and not WebGL. WebGPU has three improvements,
that from your talking are missing in WebGL?
- It has compute shaders
- It has arrays
- It has structs
- What else?[...]
On 8/1/2026 3:43 PM, Mild Shock wrote:
Hi,
WebGPU and WebGL are two different things. I explained
that towards you already like 3-5 times.
Of course we can make a special texture to handle it.
You still don't understand that I am using WebGPU,
and not WebGL. WebGPU has three improvements,
that from your talking are missing in WebGL?
- It has compute shaders
- It has arrays
- It has structs
- What else?[...]
It has textures to work with in the pipeline. But, I still don't know
what you main goal is?
On 8/1/2026 3:43 PM, Mild Shock wrote:
Hi,
WebGPU and WebGL are two different things. I explained
that towards you already like 3-5 times.
Of course we can make a special texture to handle it.
You still don't understand that I am using WebGPU,
and not WebGL. WebGPU has three improvements,
that from your talking are missing in WebGL?
- It has compute shaders
- It has arrays
- It has structs
- What else?[...]
It has textures to work with in the pipeline. But, I still don't know
what you main goal is?
shave off some of the TOPS to do Prolog inferencing
Hi,
Why would I use text inside my compute shader.
Could you tell me. The Hack VM doesn't do
textures. You are confused. There is nothing
about textures here:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
You can read the text , it says nowhere
consume or produce textures. Its not a rendering
application. I use the compute shader to run Prolog:
"At the end of 2025 we acquired a couple of
AI Laptops , that were still cheap, since
RAM prices had not yet rocketed. The intend
was to tap into the Copilot+ certified hardware,
and shave off some of the TOPS to do Prolog
inferencing. Amazingly our -C-WAM can
churn 11.4 GIGA LIPS.
GPUs have evolved form lock-step to independent
thread scheduling. This made it possible to
port the Hack VM variant, that forms the basis
for our -C-WAM, to WebGPU computer shaders.
Using NUM_SHADERS = 4096 we could produce
11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M."
Bye
Chris M. Thomasson schrieb:
On 8/1/2026 3:43 PM, Mild Shock wrote:
Hi,
WebGPU and WebGL are two different things. I explained
that towards you already like 3-5 times.
Of course we can make a special texture to handle it.
You still don't understand that I am using WebGPU,
and not WebGL. WebGPU has three improvements,
that from your talking are missing in WebGL?
- It has compute shaders
- It has arrays
- It has structs
- What else?[...]
It has textures to work with in the pipeline. But, I still don't know
what you main goal is?
Hi,
It explicity says "Prolog inferencing" in
this phrase:
shave off some of the TOPS to do Prolog inferencing
It nowhere says draw some fancy stuff into
a Web canvas.
:- multifile(strings/3).
/* de = ISO locale atoms with prefix de_ */ strings('evaluation_error.zero_divisor', de, 'Nulldivision.').
/* '' = fall back ISO locale atoms */ strings('evaluation_error.zero_divisor', '', 'Division by zero.').
Hi,
It explicity says "Prolog inferencing" in
this phrase:
shave off some of the TOPS to do Prolog inferencing
It nowhere says draw some fancy stuff into
a Web canvas.
Bye
Mild Shock schrieb:
Hi,
Why would I use text inside my compute shader.
Could you tell me. The Hack VM doesn't do
textures. You are confused. There is nothing
about textures here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
You can read the text , it says nowhere
consume or produce textures. Its not a rendering
application. I use the compute shader to run Prolog:
"At the end of 2025 we acquired a couple of
AI Laptops , that were still cheap, since
RAM prices had not yet rocketed. The intend
was to tap into the Copilot+ certified hardware,
and shave off some of the TOPS to do Prolog
inferencing. Amazingly our -C-WAM can
churn 11.4 GIGA LIPS.
GPUs have evolved form lock-step to independent
thread scheduling. This made it possible to
port the Hack VM variant, that forms the basis
for our -C-WAM, to WebGPU computer shaders.
Using NUM_SHADERS = 4096 we could produce
11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M."
Bye
Chris M. Thomasson schrieb:
On 8/1/2026 3:43 PM, Mild Shock wrote:
Hi,
WebGPU and WebGL are two different things. I explained
that towards you already like 3-5 times.
Of course we can make a special texture to handle it.
You still don't understand that I am using WebGPU,
and not WebGL. WebGPU has three improvements,
that from your talking are missing in WebGL?
- It has compute shaders
- It has arrays
- It has structs
- What else?[...]
It has textures to work with in the pipeline. But, I still don't know
what you main goal is?
| Sysop: | Amessyroom |
|---|---|
| Location: | Fayetteville, NC |
| Users: | 74 |
| Nodes: | 6 (0 / 6) |
| Uptime: | 45:26:04 |
| Calls: | 1,100 |
| Files: | 1,339 |
| Messages: | 275,372 |