Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Hi,
This is quite fun, how some TLA+ guy fears
the full state of queue like the devil in
itself. But I guess if a service rate is
low and the producer has not much to do to
produce its work items, the arrival rate
has nevertheless to adapt, and dealing
with "full states", which are wrongly
called deadlock here, is the normal:
Tutorial-style talk - BlockingQueue https://github.com/lemmy/BlockingQueue/tree/main
Prolog is in good position. The bird box
model has a redo port. So sometimes switching
from push to pull, can help without doing
Deadlock Exorcism. You can also translate
the bird box ports into pi-calculus:
A pi-calculus Specification of Prolog
Benjamin Z. Li - University of Pennsylvania
11 Apr 1994, European Symposium on Programming,
Prolog, Unification, Backtracking https://scispace.com/pdf/a-pi-calculus-specification-of-prolog-3qf2pf04ud.pdf
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Mild Shock schrieb:
Hi,
This is quite fun, how some TLA+ guy fears
the full state of queue like the devil in
itself. But I guess if a service rate is
low and the producer has not much to do to
produce its work items, the arrival rate
has nevertheless to adapt, and dealing
with "full states", which are wrongly
called deadlock here, is the normal:
Tutorial-style talk - BlockingQueue
https://github.com/lemmy/BlockingQueue/tree/main
Prolog is in good position. The bird box
model has a redo port. So sometimes switching
from push to pull, can help without doing
Deadlock Exorcism. You can also translate
the bird box ports into pi-calculus:
A pi-calculus Specification of Prolog
Benjamin Z. Li - University of Pennsylvania
11 Apr 1994, European Symposium on Programming,
Prolog, Unification, Backtracking
https://scispace.com/pdf/a-pi-calculus-specification-of-prolog-3qf2pf04ud.pdf
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Is a trivial control construct for(),
when used in a compute shader with
NVIDIA Volta evolution, i.e. MIMD,
can lead to different time spend by
individual compute shaders:
fn main(global_id : i32) {
-a-a i : i32 = 0;
-a-a while (i < globa_id) {
-a-a-a-a-a i++;
-a-a }
}
You can visiualize as the time spent
by each logical thread as follows:
global id, logical thread life line
1-a-a-a-a [-a-a-a ]
2-a-a-a-a [-a-a-a-a-a-a-a ]
3-a-a-a-a [-a-a-a-a-a-a-a-a-a-a-a ]
4-a-a-a-a [-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a ]
5-a-a-a-a [-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a ]
Etc..
With work items and load balancing you
could run the above with a lower number
of logical threads, I am writing the
work item number now inside the sub life
line inside the overall life line of
the logical thread:
worker , worker work items
A-a-a-a-a [3-a-a-a-a-a-a-a-a-a-a ]
B-a-a-a-a [4-a-a-a-a-a-a-a-a-a-a-a-a-a-a ][2-a-a-a-a-a-a ]
C-a-a-a-a [5-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a ][1-a-a ]
The overall time slightly increased by 1,
i.e. the case global_id = k combined
with the case global_id = n-k+1 . Also
one worker didn't have two work items,
only one work item. But the number of
logical threads needed was halfed.
Ok, a mpmc queue will be not that
intelligent, concerning the work sheduling.
But one could experiment with mpmc queue
priority queues etc.. etc..
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Mild Shock schrieb:
Hi,
This is quite fun, how some TLA+ guy fears
the full state of queue like the devil in
itself. But I guess if a service rate is
low and the producer has not much to do to
produce its work items, the arrival rate
has nevertheless to adapt, and dealing
with "full states", which are wrongly
called deadlock here, is the normal:
Tutorial-style talk - BlockingQueue
https://github.com/lemmy/BlockingQueue/tree/main
Prolog is in good position. The bird box
model has a redo port. So sometimes switching
from push to pull, can help without doing
Deadlock Exorcism. You can also translate
the bird box ports into pi-calculus:
A pi-calculus Specification of Prolog
Benjamin Z. Li - University of Pennsylvania
11 Apr 1994, European Symposium on Programming,
Prolog, Unification, Backtracking
https://scispace.com/pdf/a-pi-calculus-specification-of-prolog-3qf2pf04ud.pdf
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
For graphics, rendering, in the worst case your
FPS can go down. Then you have a 2x as performant
graphic card, and sundently the FPS is ok again!
Or you render into a smaller screen, with less
number of pixels, and things turn good again.
So for my pixel phone AI experiment, or what I
will toy around on a AI laptop, I am singing:
-a I'm a spinner, I'm a sinner
-a I spin on CAS loops for my dinner
-a Some call it busy-wait, I call it fate
-a When the queue is empty, I just rotate
Bye
Whats better Steve Miller or Muddy Waters?
The Joker
https://www.youtube.com/watch?v=dV3AziKTBUo
Hoochie Coochie Man
https://www.youtube.com/watch?v=e_l6A7krjrQ
Mild Shock schrieb:
Hi,
Is a trivial control construct for(),
when used in a compute shader with
NVIDIA Volta evolution, i.e. MIMD,
can lead to different time spend by
individual compute shaders:
fn main(global_id : i32) {
-a-a-a i : i32 = 0;
-a-a-a while (i < globa_id) {
-a-a-a-a-a-a i++;
-a-a-a }
}
You can visiualize as the time spent
by each logical thread as follows:
global id, logical thread life line
1-a-a-a-a [-a-a-a ]
2-a-a-a-a [-a-a-a-a-a-a-a ]
3-a-a-a-a [-a-a-a-a-a-a-a-a-a-a-a ]
4-a-a-a-a [-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a ]
5-a-a-a-a [-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a ]
Etc..
With work items and load balancing you
could run the above with a lower number
of logical threads, I am writing the
work item number now inside the sub life
line inside the overall life line of
the logical thread:
worker , worker work items
A-a-a-a-a [3-a-a-a-a-a-a-a-a-a-a ]
B-a-a-a-a [4-a-a-a-a-a-a-a-a-a-a-a-a-a-a ][2-a-a-a-a-a-a ]
C-a-a-a-a [5-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a ][1-a-a ]
The overall time slightly increased by 1,
i.e. the case global_id = k combined
with the case global_id = n-k+1 . Also
one worker didn't have two work items,
only one work item. But the number of
logical threads needed was halfed.
Ok, a mpmc queue will be not that
intelligent, concerning the work sheduling.
But one could experiment with mpmc queue
priority queues etc.. etc..
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow
https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels
https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Mild Shock schrieb:
Hi,
This is quite fun, how some TLA+ guy fears
the full state of queue like the devil in
itself. But I guess if a service rate is
low and the producer has not much to do to
produce its work items, the arrival rate
has nevertheless to adapt, and dealing
with "full states", which are wrongly
called deadlock here, is the normal:
Tutorial-style talk - BlockingQueue
https://github.com/lemmy/BlockingQueue/tree/main
Prolog is in good position. The bird box
model has a redo port. So sometimes switching
from push to pull, can help without doing
Deadlock Exorcism. You can also translate
the bird box ports into pi-calculus:
A pi-calculus Specification of Prolog
Benjamin Z. Li - University of Pennsylvania
11 Apr 1994, European Symposium on Programming,
Prolog, Unification, Backtracking
https://scispace.com/pdf/a-pi-calculus-specification-of-prolog-3qf2pf04ud.pdf
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
While HBM and RDMA happen outside of a the main
silicon chip. Amazing things are now happening
inside a silicon chip as found in AI laptops.
Basically XILINX later acquired by AMD, had
already the Versal architecture. Where FGPA was
used to custom wire chips. The Versal area
had already Network-on-Chip (NoC): https://www.adiuvoengineering.com/post/microzed-chronicles-versal-part-two-device-architecture
While a Ryzen AI 7 350 /w Radeon 860M does not
really have a versal area anymore. But the
Network-on-Chip (NoC) survived, with twist:
GEMM Performance Generations of Ryzen AI NPUs
4.3 On-The-Fly Tensor Transformations
We extensively exploit the multi-dimensional
addressing feature of DMAs to reorganize data into
tiled layouts, as needed by the NPU cores.
https://arxiv.org/abs/2512.13282v1
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Mild Shock wrote:--- Synchronet 3.22a-Linux NewsLink 1.2
Hi,
How it started:
Captain: Throw the switch, Scotty!
Enterprise: Cloaking Device makes it invisible
You stupid ass. You posted this twice.
Mild Shock wrote:
So a few days later comes out the LLaMA, I do some calculations and I
figure out rCLOkay, 65 billion parameters. You probably need about 40 gigs
of RAM, with 4-bit quantization. So this can run on a MacBook. Why not
do it?rCY
you are a shame to your mother
Perplexity Increase: Quantizing to 4-bit typically increases perplexity
Reasoning & Coding: Complex reasoning chains and coding tasks suffer
Mild Shock wrote:miles away from her.
Hi,
My mother is worried that I fucked Lane W.
aka Micro Penis mother 24 hours straight.
She was screaming, basically singing all
the arias from operas that Luciano Pavarotti
usually sings. You Lane W. aka Micro Penis
should have heard it, since you
live in the basement of your mothers house.
No, actually remarkably, I don't. According to google I live 433
Strike!
See, what i said about you was spot on.
What you said about me was generic and incorrect.
You really suck, man.
Hi,
Micro penis brain is in constant hiatus.
He can even not detect a trope.
LoL
Bye
Lane W schrieb:
Mild Shock wrote:
Hi,
My mother is worried that I fucked Lane W.
aka Micro Penis mother 24 hours straight.
She was screaming, basically singing all
the arias from operas that Luciano Pavarotti
usually sings. You Lane W. aka Micro Penis
should have heard it, since you
live in the basement of your mothers house.
No, actually remarkably, I don't. According to google I live 433miles away from her.
Strike!
See, what i said about you was spot on.
What you said about me was generic and incorrect.
You really suck, man.
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Mild Shock has no idea who I am or what I represent
Hi,
If any of you guys do not understand what
is meant by or what the implications are:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Well I wouldn't care less. There are two
outcomes for numb nuts:
- Ignoramus: They don't understand it, but
-a they will understand it before they die.
- Ignorabimus: They don't understand it, and
-a will never understand it, and they die.
So who cares, its not my problem, you people
are stupid as fuck, and slow as fuck...
Bye
Mild Shock schrieb:
Hi,
Micro penis brain is in constant hiatus.
He can even not detect a trope.
LoL
Bye
Lane W schrieb:
Mild Shock wrote:miles away from her.
Hi,
My mother is worried that I fucked Lane W.
aka Micro Penis mother 24 hours straight.
She was screaming, basically singing all
the arias from operas that Luciano Pavarotti
usually sings. You Lane W. aka Micro Penis
should have heard it, since you
live in the basement of your mothers house.
No, actually remarkably, I don't. According to google I live 433
Strike!
See, what i said about you was spot on.
What you said about me was generic and incorrect.
You really suck, man.
Hi,
Now you can compare this here from 2008
with modern AI Laptops for 500-1000 USD:
Google spotlights data center inner workings https://web.archive.org/web/20131019063218/http://news.cnet.com/8301-10784_3-9955184-7.html
There is a striking similarity, only what
once occupied a rack, has now the size
of your plam, all inside one silicon chip:
- Multiple CPU cores on the same chip
- Multiple GPU units on the same chip
- Network on the same chip communication
- Crossbar caches on the same chip
- Disk controllers on the same chip
- Multi channel RAM access on the same chip
Pretty cool!
Bye
P.S.: Example such devices with iGPU:
Intel(R) Core(TM) Ultra 7 258V
AMD Ryzen AI 7 350 w/ Radeon 860M
Apple A18 Pro, Darwin Kernel Version 25.5.0
Snapdragon(R) X - X126100 - Qualcomm(R) Oryon(TM) CPU
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
From: Mild Shock <janburse@fastmail.fm>
Subject: NVIDIA evacuated its Chinese market [Tau Scaling]
Date: Thu, 23 Jul 2026 19:13:51 +0200
How it started:
Captain: Throw the switch, Scotty!
Enterprise: Cloaking Device makes it invisible
Spock: Military secrets are the most fleeting of all.
Kirk Escapes the Romulans - The Enterprise Incident https://www.youtube.com/watch?v=AusAGjwlql8
Hi,
You are a moron, and you represent putin payed
trolls from the army of brainless troll morons.
Bye
Lane W schrieb:
Mild Shock has no idea who I am or what I represent
Mild Shock schrieb:
Hi,
If any of you guys do not understand what
is meant by or what the implications are:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Well I wouldn't care less. There are two
outcomes for numb nuts:
- Ignoramus: They don't understand it, but
-a-a they will understand it before they die.
- Ignorabimus: They don't understand it, and
-a-a will never understand it, and they die.
So who cares, its not my problem, you people
are stupid as fuck, and slow as fuck...
Bye
Mild Shock schrieb:
Hi,
Micro penis brain is in constant hiatus.
He can even not detect a trope.
LoL
Bye
Lane W schrieb:
Mild Shock wrote:miles away from her.
Hi,
My mother is worried that I fucked Lane W.
aka Micro Penis mother 24 hours straight.
She was screaming, basically singing all
the arias from operas that Luciano Pavarotti
usually sings. You Lane W. aka Micro Penis
should have heard it, since you
live in the basement of your mothers house.
No, actually remarkably, I don't. According to google I live 433
Strike!
See, what i said about you was spot on.
What you said about me was generic and incorrect.
You really suck, man.
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Show an outline of what you
need you compute shader to do?
Hi,
If any of you guys do not understand what
is meant by or what the implications are:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Well I wouldn't care less. There are two
outcomes for numb nuts:
- Ignoramus: They don't understand it, but
-a they will understand it before they die.
- Ignorabimus: They don't understand it, and
-a will never understand it, and they die.
So who cares, its not my problem, you people
are stupid as fuck, and slow as fuck...
Bye
Mild Shock schrieb:
Hi,
Micro penis brain is in constant hiatus.
He can even not detect a trope.
LoL
Bye
Lane W schrieb:
Mild Shock wrote:miles away from her.
Hi,
My mother is worried that I fucked Lane W.
aka Micro Penis mother 24 hours straight.
She was screaming, basically singing all
the arias from operas that Luciano Pavarotti
usually sings. You Lane W. aka Micro Penis
should have heard it, since you
live in the basement of your mothers house.
No, actually remarkably, I don't. According to google I live 433
Strike!
See, what i said about you was spot on.
What you said about me was generic and incorrect.
You really suck, man.
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
the first product was not announced until May 2017 https://en.wikipedia.org/wiki/Volta_%28microarchitecture%29
Intel(R) Core(TM) Ultra 7 258V
AMD Ryzen AI 7 350 w/ Radeon 860M
Apple A18 Pro, Darwin Kernel Version 25.5.0
Snapdragon(R) X - X126100 - Qualcomm(R) Oryon(TM) CPU
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979} https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
A Sputnik Commodore C64 with 8088
from the basement of your mother
11.4 Giga Lips with a Budget Laptop
At the end of 2025 we acquired a couple of AI Laptops https://github.com/Jean-Luc-Picard-2021/gigabudget
Hi,
Its not tested on some Single Instruction/
Multiple Data (SIMD) GPU. It was only tested on
AI Laptops with Multiple instruction, Multiple
Data (GPU) architecture for the scalar registers
per logical thread. As introduced by NVIDIA Volta
in around 2017:
the first product was not announced until May 2017 https://en.wikipedia.org/wiki/Volta_%28microarchitecture%29
Although I wrote the code of Hack VM with SIMD
in mind, I never tested it on a pure SIMD GPU,
and I never ported boot.mjs or boot2.mjs to
WebGL2 / GLSL. I uploaded WebGPU / WGSL. Among the
tester I had were these AI Laptops, that could all
run WebGPU / WGSL in a browser:
Intel(R) Core(TM) Ultra 7 258V
AMD Ryzen AI 7 350 w/ Radeon 860M
Apple A18 Pro, Darwin Kernel Version 25.5.0
Snapdragon(R) X - X126100 - Qualcomm(R) Oryon(TM) CPU
Some AI Laptops had WebGPU / WGSL still behind
a browser flag, since its relatively new on ARM.
Also the above AI Laptops have all a iGPU and
not a separate GPU card.
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Because I use WebGPU and not WebGL. And
because WebGPU can adresss modern GPU
developed with the NVIDIA Volta evolution,
which happened in 2017. Namley that compute
shaders are not any more subject to the
realization restriction of lock step
execution, but have independent thread state.
And because there is independent thread state
there is also independent time spent for a
a work item by each logical thread, if the
submitted logical thread uses a lot of branching
logic or even loops. But the use of branching
and loops is encouraged in independent thread
state programming of compute shaders. The variables
that can drive such logic are the scalar variables:
Tour of WGSL - Control Flow https://google.github.io/tour-of-wgsl/control-flow/
Then not to waste GPU compute time, by logical
threads doing nothing. You will need to
introduce some load balancing among multiple
logical threads. And MPMC queues are one way to
readize load balancing. Compute shaders with
producer and consumer entry points are proposed
as fundamental architecture by Thunder Kittens:
ThunderKittens: Simple, Fast, and Adorable AI Kernels https://arxiv.org/abs/2410.20399
They are used by this SpaceX acquisition:
Composer 2 Technical Report
https://arxiv.org/abs/2603.24477
Thunder Kittens uses Hardware support, i.e. tma_expect().
Bye
Chris M. Thomasson schrieb:
never meant to be used in a GPU.
Dmitry CAS version can be used, but
Why do you even need a mpmc queue
in your compute shader anyway?
Mild Shock schrieb:
Hi,
This is quite fun, how some TLA+ guy fears
the full state of queue like the devil in
itself. But I guess if a service rate is
low and the producer has not much to do to
produce its work items, the arrival rate
has nevertheless to adapt, and dealing
with "full states", which are wrongly
called deadlock here, is the normal:
Tutorial-style talk - BlockingQueue
https://github.com/lemmy/BlockingQueue/tree/main
Prolog is in good position. The bird box
model has a redo port. So sometimes switching
from push to pull, can help without doing
Deadlock Exorcism. You can also translate
the bird box ports into pi-calculus:
A pi-calculus Specification of Prolog
Benjamin Z. Li - University of Pennsylvania
11 Apr 1994, European Symposium on Programming,
Prolog, Unification, Backtracking
https://scispace.com/pdf/a-pi-calculus-specification-of-prolog-3qf2pf04ud.pdf
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Its not tested on some Single Instruction/
Multiple Data (SIMD) GPU. It was only tested on
AI Laptops with Multiple instruction, Multiple
Data (GPU) architecture for the scalar registers
per logical thread. As introduced by NVIDIA Volta
in around 2017:
the first product was not announced until May 2017 https://en.wikipedia.org/wiki/Volta_%28microarchitecture%29
Although I wrote the code of Hack VM with SIMD
in mind, I never tested it on a pure SIMD GPU,
and I never ported boot.mjs or boot2.mjs to
WebGL2 / GLSL. I uploaded WebGPU / WGSL. Among the
tester I had were these AI Laptops, that could all
run WebGPU / WGSL in a browser:
Intel(R) Core(TM) Ultra 7 258V
AMD Ryzen AI 7 350 w/ Radeon 860M
Apple A18 Pro, Darwin Kernel Version 25.5.0
Snapdragon(R) X - X126100 - Qualcomm(R) Oryon(TM) CPU
Some AI Laptops had WebGPU / WGSL still behind
a browser flag, since its relatively new on ARM.
Also the above AI Laptops have all a iGPU and
not a separate GPU card.
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
-a-a-a public static void run() {
-a-a-a-a-a-a-a for (int A = 1; A < 192; A++) {
-a-a-a-a-a-a-a-a-a-a-a int Y = (771-A)/3;
-a-a-a-a-a-a-a-a-a-a-a for (int B = A; B < Y; B++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int Z = (771-A-B)/2;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a for (int C = B; C < Z; C++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int D = 711-A-B-C;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a if (A*B*C == 711000000/D &&
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a 711000000 % D == 0)
-a-a-a System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a }
-a-a-a }
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Mild Shock schrieb:
Hi,
Its not tested on some Single Instruction/
Multiple Data (SIMD) GPU. It was only tested on
AI Laptops with Multiple instruction, Multiple
Data (GPU) architecture for the scalar registers
per logical thread. As introduced by NVIDIA Volta
in around 2017:
the first product was not announced until May 2017
https://en.wikipedia.org/wiki/Volta_%28microarchitecture%29
Although I wrote the code of Hack VM with SIMD
in mind, I never tested it on a pure SIMD GPU,
and I never ported boot.mjs or boot2.mjs to
WebGL2 / GLSL. I uploaded WebGPU / WGSL. Among the
tester I had were these AI Laptops, that could all
run WebGPU / WGSL in a browser:
Intel(R) Core(TM) Ultra 7 258V
AMD Ryzen AI 7 350 w/ Radeon 860M
Apple A18 Pro, Darwin Kernel Version 25.5.0
Snapdragon(R) X - X126100 - Qualcomm(R) Oryon(TM) CPU
Some AI Laptops had WebGPU / WGSL still behind
a browser flag, since its relatively new on ARM.
Also the above AI Laptops have all a iGPU and
not a separate GPU card.
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
public static void run() {
for (int A = 1; A < 192; A++) {
int Y = (771-A)/3;
for (int B = A; B < Y; B++) {
int Z = (771-A-B)/2;
for (int C = B; C < Z; C++) {
int D = 711-A-B-C;
if (A*B*C == 711000000/D &&
711000000 % D == 0)
System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
}
}
}
}
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
public static void run() {
for (int A = 1; A < 192; A++) {
int Y = (771-A)/3;
for (int B = A; B < Y; B++) {
int Z = (771-A-B)/2;
for (int C = B; C < Z; C++) {
int D = 711-A-B-C;
if (A*B*C == 711000000/D &&
711000000 % D == 0)
System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
}
}
}
}
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Hi,
Whats this "forget" trope of glue sniffing
Rossy Boy with his herpes blisters?
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Why should I forget Bulgarians,
they are never on my mind. Do you
see me doing ggml stuff?
I only hypothesized that it is
over for Python as the machine
learning language or AI inferencing
locally on AI laptops language, and
made the ggml case, so I already forgot
about them. Which might give you a glimps,
why WebGPU was used for this here:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Is an interesting choice. Even
github has some Languages statistics,
giving an account what I used:
HTML 67.5% JavaScript 23.1% CSS 9.4%
Have Fun!
Bye
P.S.: The example below is not p-adics,
you complete imbecil moron. Its just:
7-11 cubic Solution by Pritchard & Gries https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
Ross Finlayson schrieb:
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
-a-a-a-a-a public static void run() {
-a-a-a-a-a-a-a-a-a for (int A = 1; A < 192; A++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a int Y = (771-A)/3;
-a-a-a-a-a-a-a-a-a-a-a-a-a for (int B = A; B < Y; B++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int Z = (771-A-B)/2;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a for (int C = B; C < Z; C++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int D = 711-A-B-C;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a if (A*B*C == 711000000/D &&
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a 711000000 % D == 0)
-a-a-a-a-a System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a }
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Hi,
Some counter PyTorch Python trends are
for example OpenAIs Triton. And the variant
miniTriton CUDA vibe produced by Kimi K3 (sic!):
"We further tested whether Kimi K3 could build
a GPU programming system from scratch. Kimi K3
developed MiniTriton, a compact Triton-like
compiler with its own tile-level IR layer over
MLIR, optimization passes, and a PTX code-
generation pipeline.
Across supported roofline benchmarks, MiniTriton
delivers performance on par with or better than
Triton and torch.compile rCo beating Triton on
certain workloads. Beyond microbenchmarks,
MiniTriton sustains end-to-end nanoGPT training
with stable convergence, the loss curve
closely tracking the reference with only minor
divergence rCo validating the full pipeline on a
realistic workload. These results demonstrate
that Kimi K3 can build a coherent end-to-end
compiler rCo from DSL frontend and IR passes to
PTX codegen and runtime rCo rather than isolated
kernels; its from-scratch Tensor Core path
already rivals TritonrCOs extensively optimized stack."
GPU Compiler Development
https://www.kimi.com/blog/kimi-k3
Although many GPU corporate stuff is anonymized,
and some AI papers have lists of 30 authors. Here
nanoGPT is mentioned which is tied to the name
Andrej Karpathy. See also here:
Update Nov 2025 nanoGPT has a new and
improved cousin called nanochat.
https://github.com/karpathy/nanogpt
But as can be seen, he moved on to another project.
Bye
Mild Shock schrieb:
Hi,
Whats this "forget" trope of glue sniffing
Rossy Boy with his herpes blisters?
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Why should I forget Bulgarians,
they are never on my mind. Do you
see me doing ggml stuff?
I only hypothesized that it is
over for Python as the machine
learning language or AI inferencing
locally on AI laptops language, and
made the ggml case, so I already forgot
about them. Which might give you a glimps,
why WebGPU was used for this here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Is an interesting choice. Even
github has some Languages statistics,
giving an account what I used:
HTML 67.5% JavaScript 23.1% CSS 9.4%
Have Fun!
Bye
P.S.: The example below is not p-adics,
you complete imbecil moron. Its just:
7-11 cubic Solution by Pritchard & Gries
https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
Ross Finlayson schrieb:
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
-a-a-a-a-a public static void run() {
-a-a-a-a-a-a-a-a-a for (int A = 1; A < 192; A++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a int Y = (771-A)/3;
-a-a-a-a-a-a-a-a-a-a-a-a-a for (int B = A; B < Y; B++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int Z = (771-A-B)/2;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a for (int C = B; C < Z; C++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int D = 711-A-B-C;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a if (A*B*C == 711000000/D && >> -a>>-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a 711000000 % D == 0)
-a-a-a-a-a System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a }
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Hi,
Whats this "forget" trope of glue sniffing
Rossy Boy with his herpes blisters?
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Why should I forget Bulgarians,
they are never on my mind. Do you
see me doing ggml stuff?
I only hypothesized that it is
over for Python as the machine
learning language or AI inferencing
locally on AI laptops language, and
made the ggml case, so I already forgot
about them. Which might give you a glimps,
why WebGPU was used for this here:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Is an interesting choice. Even
github has some Languages statistics,
giving an account what I used:
HTML 67.5% JavaScript 23.1% CSS 9.4%
Have Fun!
Bye
P.S.: The example below is not p-adics,
you complete imbecil moron. Its just:
7-11 cubic Solution by Pritchard & Gries https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
Ross Finlayson schrieb:
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
-a-a-a-a-a public static void run() {
-a-a-a-a-a-a-a-a-a for (int A = 1; A < 192; A++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a int Y = (771-A)/3;
-a-a-a-a-a-a-a-a-a-a-a-a-a for (int B = A; B < Y; B++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int Z = (771-A-B)/2;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a for (int C = B; C < Z; C++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int D = 711-A-B-C;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a if (A*B*C == 711000000/D &&
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a 711000000 % D == 0)
-a-a-a-a-a System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a }
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Hi,
While Huggingfaces hired GG in 2026,
AK was hired by Anthropic in 2026:
Andrej Karpathy (born 23 October 1986[3])
is a Slovak-Canadian AI researcher, who
co-founded and formerly worked at OpenAI
In 2026 he joined Anthropic as part of
the pretraining team.
https://en.wikipedia.org/wiki/Andrej_Karpathy
But his nanochat archivement has an
interesting time line:
168 hours , Original OpenAI GPT-2 checkpoint, 2019
3 hours , d24 baseline, slightly overtrained, Jan 29 2026
1 1/2 hour, autoresearch round 2, Mar 14 2026
The best ChatGPT that $100 can buy.
https://github.com/karpathy/nanochat
But what hardware was the enabler. What is the
NVIDIA H100 GPU even. Well the thingy is surely not
a Budget Laptop, performance pretty much
dependence on data elememt size, the H100 NVL
version (*), and when using tensor operations,
and not only scalar operations:
8-bit towards 3000 tera flops
16-bit towards 1500 tera flops
32-bit towards 900 tera flops
Cool! I guess this experiment would tap into 60
tera flops, since it only uses scalar operations so far:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
You could perform it by migration the web application
using WebGPU into a node.js standalone application
using the dawn library for GPU access.
Bye
(*) https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet
Mild Shock schrieb:
Hi,
Whats this "forget" trope of glue sniffing
Rossy Boy with his herpes blisters?
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Why should I forget Bulgarians,
they are never on my mind. Do you
see me doing ggml stuff?
I only hypothesized that it is
over for Python as the machine
learning language or AI inferencing
locally on AI laptops language, and
made the ggml case, so I already forgot
about them. Which might give you a glimps,
why WebGPU was used for this here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Is an interesting choice. Even
github has some Languages statistics,
giving an account what I used:
HTML 67.5% JavaScript 23.1% CSS 9.4%
Have Fun!
Bye
P.S.: The example below is not p-adics,
you complete imbecil moron. Its just:
7-11 cubic Solution by Pritchard & Gries
https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
Ross Finlayson schrieb:
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
-a-a-a-a-a public static void run() {
-a-a-a-a-a-a-a-a-a for (int A = 1; A < 192; A++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a int Y = (771-A)/3;
-a-a-a-a-a-a-a-a-a-a-a-a-a for (int B = A; B < Y; B++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int Z = (771-A-B)/2;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a for (int C = B; C < Z; C++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int D = 711-A-B-C;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a if (A*B*C == 711000000/D && >> -a>>-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a 711000000 % D == 0)
-a-a-a-a-a System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a }
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Hi,
One could critisize that my -C-WAM doesn't
utilize GPU to the fullest, since its GPU
backend prototype only uses scalar operations
and no vector or matrix operations. And
modern GPUs thrive on vector and matrix
operations. Especially matrix operations giving
a boost of a factor 15x or so. There are
many papers already showing how Prolog can be
mapped to matrix operations. Only this research
is completely ignored by Prolog systems such as
SICStus, Ciao, SWI, ECLiPSe etc.. But lets
illustrate what vector operations could do
for -C-WAM, take this compilation of the Prolog
goal between(0,1023,X), Y is X*2+3:
int X;
int Y;
for (X=0; X < 1024; X++) {
-a-a-a Y=X*2+3;
-a-a-a [...]
}
With vector operations, and vectors of size
32 one could do:
int X1;
int[] X = new int[32];
int X3;
int[] Y = new int[32];
for (X1 = 0; X1 < 1024 / 32; X1++) {
-a-a-a for (int X2 = 0; X2 < 32; X2++)
-a-a-a-a-a-a X[X2] = X1*32+X2;
-a-a-a vec_mul_add(X, 2, 3, Y);
-a-a-a [..]
}
Have Fun!
Bye
Mild Shock schrieb:
Hi,
While Huggingfaces hired GG in 2026,
AK was hired by Anthropic in 2026:
Andrej Karpathy (born 23 October 1986[3])
is a Slovak-Canadian AI researcher, who
co-founded and formerly worked at OpenAI
In 2026 he joined Anthropic as part of
the pretraining team.
https://en.wikipedia.org/wiki/Andrej_Karpathy
But his nanochat archivement has an
interesting time line:
168 hours , Original OpenAI GPT-2 checkpoint, 2019
3 hours , d24 baseline, slightly overtrained, Jan 29 2026
1 1/2 hour, autoresearch round 2, Mar 14 2026
The best ChatGPT that $100 can buy.
https://github.com/karpathy/nanochat
But what hardware was the enabler. What is the
NVIDIA H100 GPU even. Well the thingy is surely not
a Budget Laptop, performance pretty much
dependence on data elememt size, the H100 NVL
version (*), and when using tensor operations,
and not only scalar operations:
8-bit towards 3000 tera flops
16-bit towards 1500 tera flops
32-bit towards 900 tera flops
Cool! I guess this experiment would tap into 60
tera flops, since it only uses scalar operations so far:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
You could perform it by migration the web application
using WebGPU into a node.js standalone application
using the dawn library for GPU access.
Bye
(*)
https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet
Mild Shock schrieb:
Hi,
Whats this "forget" trope of glue sniffing
Rossy Boy with his herpes blisters?
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Why should I forget Bulgarians,
they are never on my mind. Do you
see me doing ggml stuff?
I only hypothesized that it is
over for Python as the machine
learning language or AI inferencing
locally on AI laptops language, and
made the ggml case, so I already forgot
about them. Which might give you a glimps,
why WebGPU was used for this here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Is an interesting choice. Even
github has some Languages statistics,
giving an account what I used:
HTML 67.5% JavaScript 23.1% CSS 9.4%
Have Fun!
Bye
P.S.: The example below is not p-adics,
you complete imbecil moron. Its just:
7-11 cubic Solution by Pritchard & Gries
https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
Ross Finlayson schrieb:
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
-a-a-a-a-a public static void run() {
-a-a-a-a-a-a-a-a-a for (int A = 1; A < 192; A++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a int Y = (771-A)/3;
-a-a-a-a-a-a-a-a-a-a-a-a-a for (int B = A; B < Y; B++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int Z = (771-A-B)/2;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a for (int C = B; C < Z; C++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int D = 711-A-B-C;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a if (A*B*C == 711000000/D && >>> -a>>-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a 711000000 % D == 0)
-a-a-a-a-a System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a }
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Hi,
The nice thing about AI accelerators, pioneered
maybe by Apple Silicon and their unified memory.
The AMD APU model can be extended so that
vector and matrix operations become uniformly
available for GPU and CPU. With unified memory
already a vector operation such as:
vec_mul_add(X, 2, 3, Y)
Only needs the X and Y address. But I havent
got my head around yet how this is all organized.
Maybe a GPU has still its own GEMM cores,
but you find Apple Silicon C++/C source code,
that taps into vector and matrix operations
by Zero Copying. The Copying is left to the DMA
of the vector or matrix operation. And moderated
by the various caches. Leading to the slogan, that
multiple floating point operations become zero cost:
Some teaching can be found here https://www.hpc-ch.org/category/topics/course-workshop/
Bye
Mild Shock schrieb:
Hi,
One could critisize that my -C-WAM doesn't
utilize GPU to the fullest, since its GPU
backend prototype only uses scalar operations
and no vector or matrix operations. And
modern GPUs thrive on vector and matrix
operations. Especially matrix operations giving
a boost of a factor 15x or so. There are
many papers already showing how Prolog can be
mapped to matrix operations. Only this research
is completely ignored by Prolog systems such as
SICStus, Ciao, SWI, ECLiPSe etc.. But lets
illustrate what vector operations could do
for -C-WAM, take this compilation of the Prolog
goal between(0,1023,X), Y is X*2+3:
int X;
int Y;
for (X=0; X < 1024; X++) {
-a-a-a-a Y=X*2+3;
-a-a-a-a [...]
}
With vector operations, and vectors of size
32 one could do:
int X1;
int[] X = new int[32];
int X3;
int[] Y = new int[32];
for (X1 = 0; X1 < 1024 / 32; X1++) {
-a-a-a-a for (int X2 = 0; X2 < 32; X2++)
-a-a-a-a-a-a-a X[X2] = X1*32+X2;
-a-a-a-a vec_mul_add(X, 2, 3, Y);
-a-a-a-a [..]
}
Have Fun!
Bye
Mild Shock schrieb:
Hi,
While Huggingfaces hired GG in 2026,
AK was hired by Anthropic in 2026:
Andrej Karpathy (born 23 October 1986[3])
is a Slovak-Canadian AI researcher, who
co-founded and formerly worked at OpenAI
In 2026 he joined Anthropic as part of
the pretraining team.
https://en.wikipedia.org/wiki/Andrej_Karpathy
But his nanochat archivement has an
interesting time line:
168 hours , Original OpenAI GPT-2 checkpoint, 2019
3 hours , d24 baseline, slightly overtrained, Jan 29 2026
1 1/2 hour, autoresearch round 2, Mar 14 2026
The best ChatGPT that $100 can buy.
https://github.com/karpathy/nanochat
But what hardware was the enabler. What is the
NVIDIA H100 GPU even. Well the thingy is surely not
a Budget Laptop, performance pretty much
dependence on data elememt size, the H100 NVL
version (*), and when using tensor operations,
and not only scalar operations:
8-bit towards 3000 tera flops
16-bit towards 1500 tera flops
32-bit towards 900 tera flops
Cool! I guess this experiment would tap into 60
tera flops, since it only uses scalar operations so far:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
You could perform it by migration the web application
using WebGPU into a node.js standalone application
using the dawn library for GPU access.
Bye
(*)
https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet
Mild Shock schrieb:
Hi,
Whats this "forget" trope of glue sniffing
Rossy Boy with his herpes blisters?
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Why should I forget Bulgarians,
they are never on my mind. Do you
see me doing ggml stuff?
I only hypothesized that it is
over for Python as the machine
learning language or AI inferencing
locally on AI laptops language, and
made the ggml case, so I already forgot
about them. Which might give you a glimps,
why WebGPU was used for this here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Is an interesting choice. Even
github has some Languages statistics,
giving an account what I used:
HTML 67.5% JavaScript 23.1% CSS 9.4%
Have Fun!
Bye
P.S.: The example below is not p-adics,
you complete imbecil moron. Its just:
7-11 cubic Solution by Pritchard & Gries
https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
Ross Finlayson schrieb:
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
-a-a-a-a-a public static void run() {
-a-a-a-a-a-a-a-a-a for (int A = 1; A < 192; A++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a int Y = (771-A)/3;
-a-a-a-a-a-a-a-a-a-a-a-a-a for (int B = A; B < Y; B++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int Z = (771-A-B)/2;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a for (int C = B; C < Z; C++) {
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a int D = 711-A-B-C;
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a if (A*B*C == 711000000/D && >>>> -a>>-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a 711000000 % D == 0)
-a-a-a-a-a System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a-a-a-a-a }
-a-a-a-a-a }
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Hi,
But the example gives also way to vector
and matrix registers. The int[] X and
int[] Y could be also held in vector
registers. Compilers can also optimize
away int[] Y, and use a inline modification,
in case X isn't used later, then playing
the role of Y:
vec_mul_add(X, 2, 3, X)
Vector and matrix registers in modern GPUs
emerged from distinct architectural milestones:
vector-like register files developed with
early programmable 3D vertex/pixel pipelines
in the late 1990s to early 2000s. While
dedicated multi-dimensional matrix registers
(Tensor Cores/Matrix Cores) were invented by
NVIDIA in 2017, starting with the Tesla
V100 (Volta microarchitecture):
From Volta To Blackwell https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell
You see the scheduling of tensure core occupation
scheduling in the above article, including memory
and register flow, following the section:
MMA Instruction Overview
It went through a couple of generations, leading
to Tensor Memory (TMEM) and collective operations,
basically realizing the PIM idea:
Processing-in-Memory Tutorials https://www.sigarch.org/processing-in-memory-tutorials-experiences-from-past-two-years-and-thoughts-looking-forward/
Have Fun!
Bye
Mild Shock schrieb:
Hi,
The nice thing about AI accelerators, pioneered
maybe by Apple Silicon and their unified memory.
The AMD APU model can be extended so that
vector and matrix operations become uniformly
available for GPU and CPU. With unified memory
already a vector operation such as:
vec_mul_add(X, 2, 3, Y)
Only needs the X and Y address. But I havent
got my head around yet how this is all organized.
Maybe a GPU has still its own GEMM cores,
but you find Apple Silicon C++/C source code,
that taps into vector and matrix operations
by Zero Copying. The Copying is left to the DMA
of the vector or matrix operation. And moderated
by the various caches. Leading to the slogan, that
multiple floating point operations become zero cost:
Some teaching can be found here
https://www.hpc-ch.org/category/topics/course-workshop/
Bye
Mild Shock schrieb:
Hi,
One could critisize that my -C-WAM doesn't
utilize GPU to the fullest, since its GPU
backend prototype only uses scalar operations
and no vector or matrix operations. And
modern GPUs thrive on vector and matrix
operations. Especially matrix operations giving
a boost of a factor 15x or so. There are
many papers already showing how Prolog can be
mapped to matrix operations. Only this research
is completely ignored by Prolog systems such as
SICStus, Ciao, SWI, ECLiPSe etc.. But lets
illustrate what vector operations could do
for -C-WAM, take this compilation of the Prolog
goal between(0,1023,X), Y is X*2+3:
int X;
int Y;
for (X=0; X < 1024; X++) {
Y=X*2+3;
[...]
}
With vector operations, and vectors of size
32 one could do:
int X1;
int[] X = new int[32];
int X3;
int[] Y = new int[32];
for (X1 = 0; X1 < 1024 / 32; X1++) {
for (int X2 = 0; X2 < 32; X2++)
X[X2] = X1*32+X2;
vec_mul_add(X, 2, 3, Y);
[..]
}
Have Fun!
Bye
Mild Shock schrieb:
Hi,
While Huggingfaces hired GG in 2026,
AK was hired by Anthropic in 2026:
Andrej Karpathy (born 23 October 1986[3])
is a Slovak-Canadian AI researcher, who
co-founded and formerly worked at OpenAI
In 2026 he joined Anthropic as part of
the pretraining team.
https://en.wikipedia.org/wiki/Andrej_Karpathy
But his nanochat archivement has an
interesting time line:
168 hours , Original OpenAI GPT-2 checkpoint, 2019
3 hours , d24 baseline, slightly overtrained, Jan 29 2026
1 1/2 hour, autoresearch round 2, Mar 14 2026
The best ChatGPT that $100 can buy.
https://github.com/karpathy/nanochat
But what hardware was the enabler. What is the
NVIDIA H100 GPU even. Well the thingy is surely not
a Budget Laptop, performance pretty much
dependence on data elememt size, the H100 NVL
version (*), and when using tensor operations,
and not only scalar operations:
8-bit towards 3000 tera flops
16-bit towards 1500 tera flops
32-bit towards 900 tera flops
Cool! I guess this experiment would tap into 60
tera flops, since it only uses scalar operations so far:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
You could perform it by migration the web application
using WebGPU into a node.js standalone application
using the dawn library for GPU access.
Bye
(*)
https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet
Mild Shock schrieb:
Hi,
Whats this "forget" trope of glue sniffing
Rossy Boy with his herpes blisters?
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Why should I forget Bulgarians,
they are never on my mind. Do you
see me doing ggml stuff?
I only hypothesized that it is
over for Python as the machine
learning language or AI inferencing
locally on AI laptops language, and
made the ggml case, so I already forgot
about them. Which might give you a glimps,
why WebGPU was used for this here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Is an interesting choice. Even
github has some Languages statistics,
giving an account what I used:
HTML 67.5% JavaScript 23.1% CSS 9.4%
Have Fun!
Bye
P.S.: The example below is not p-adics,
you complete imbecil moron. Its just:
7-11 cubic Solution by Pritchard & Gries
https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
Ross Finlayson schrieb:
On 07/26/2026 10:52 AM, Mild Shock wrote:
Hi,
You see it all boils down to find your inner peace
by an immaculate inception of some queue datatype.
KOAN/Fortran-S was an early 1990s research programming
system for distributed-memory multiprocessors . Developed
at ENS Lyon in the early 1990s . Often listed alongside
other historical parallel programming efforts.
The Message Passing: The research explicitly
compared the SVM approach against message passing
on the same hardware . The finding was that SVM
could achieve good performance without the low-level
complexity of managing explicit messages, though
the best results often came from a hybrid approach (sic!)
Here is an interesting baseline, from Java,
a class ElevenSingle that only does:
public static void run() {
for (int A = 1; A < 192; A++) {
int Y = (771-A)/3;
for (int B = A; B < Y; B++) {
int Z = (771-A-B)/2;
for (int C = B; C < Z; C++) {
int D = 711-A-B-C;
if (A*B*C == 711000000/D &&
711000000 % D == 0)
System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
}
}
}
}
And then compare it to ElevenMulti, doing some
Work Balancing Scheduler Tetris Game with 8 cores:
ElevenSingle
A=120, B=125, C=150, D=316
6.628 ms
ElevenMulti
A=120, B=125, C=150, D=316
1.941 ms
Not great, not terrible!
Bye
Oh, that's just "tricks of p-adic arithmetic".
Like other sock-puppet howler trolls, when confronted
with its base incredulity, it will descend to its
lower levers of the pathos variety.
You might be happier learning about Julia trees and
raster ops, instead of shilling yet another Ramanujan
series without saying how it's made.
Bulgarians, that's some real Boris and Natasha crap,
forget Hungarians and Bulgarians.
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979} https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
On 07/27/2026 04:22 AM, Mild Shock wrote:
Hi,That's bullshit, and alike those talking heads that
But the example gives also way to vector
and matrix registers. The int[] X and
int[] Y could be also held in vector
registers. Compilers can also optimize
away int[] Y, and use a inline modification,
in case X isn't used later, then playing
the role of Y:
vec_mul_add(X, 2, 3, X)
Vector and matrix registers in modern GPUs
emerged from distinct architectural milestones:
vector-like register files developed with
early programmable 3D vertex/pixel pipelines
in the late 1990s to early 2000s. While
dedicated multi-dimensional matrix registers
(Tensor Cores/Matrix Cores) were invented by
NVIDIA in 2017, starting with the Tesla
V100 (Volta microarchitecture):
-aFrom Volta To Blackwell
https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell
You see the scheduling of tensure core occupation
scheduling in the above article, including memory
and register flow, following the section:
MMA Instruction Overview
It went through a couple of generations, leading
to Tensor Memory (TMEM) and collective operations,
basically realizing the PIM idea:
Processing-in-Memory Tutorials
https://www.sigarch.org/processing-in-memory-tutorials-experiences-from-past-two-years-and-thoughts-looking-forward/
Have Fun!
Bye
sniff their way into talking about many-core jumbo-trons,
the super-scalar is as old as the scalar and Cray and examples alike
the Connection Machine what made all the craze of neural nets
is old-wrapped-as-new.
Fabless chips did it already.
Data centers should pay a 10000% excise on electricity,
wherever it comes from, a natural regulator of inverted economies.
And by ten thousand percent I really mean a ten thousand percent.
Data centers should pay a 10000% excise on electricity,
wherever it comes from, a natural regulator of inverted economies.
And by ten thousand percent I really mean a ten thousand percent.
generalizes better than it did before training and more importantly/\ less
it takes maybe an order of magnitude crunching to produce a good answer
than the usual over-fit answer.--- Synchronet 3.22a-Linux NewsLink 1.2
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
In comp.lang.prolog Ross Finlayson <ross.a.finlayson@gmail.com> wrote:
...
Data centers should pay a 10000% excise on electricity,
wherever it comes from, a natural regulator of inverted economies.
And by ten thousand percent I really mean a ten thousand percent.
And what would a huge surcharge do?
Almost always end up affecting the less powerful end of society
with increased costs to services the AI industry will be doing
more and more of over time.
I started a little data center (exaflops.com) many years ago.
In those distant days people (in fact one was a prof of computer
science) told me you could never make money running a supercomputer.
LOL. :)
I've had many years to watch the trends and a far more efficient
way to solve resource problems in this area is to change the
algorithms. There is vast room for improvement, mostly because
of prevailing attitudes.
I used to do competetion data science as a sideline. Companies
would pay almost any price to get an extra decimal place in
the accuracy of their forecasting processes. But typically
they were trying to supercharge a system that should be scrapped
and re-designed from scratch. One area I'm thinking of is
investment. I had a customer one time -- like many times --
ask to improve a system that predicted the future price of
various stocks. The idea (for them) was to have as accurate a
prediction of what some stock would be worth in a week or a month's
time so that some moron could use the information to decide when
to buy or sell the thing.
I tried to argue the efficient thing was to create a system that
takes the human out of the loop altogether. It doesnt provide info
for someone to decide whether or not to follow the advice --
that is just introducing more noise into the loop and probably
cancels any benefit of adding a couple decimal places of precision.
What you *should* do is make a system that is tuned to robustly
maximize the profit from managing a portfolio.
Of course they wouldnt come at that. You can't suggest taking the
managers out of the loop. :)
Another idea relevant to current AI methods might be to curtail
use of typical neural net algorithms. Many of them try to squeeze
the best performance of some NN during the training phase in
the hope the resulting system will generalize well enough to be useful
on new data. But there's kind-of a law that the harder you train
some system to perform a task well, the less well they can subsuently
perform a more general version of the same thing. It's amusing when
you look at the graphs of NN being trained and then tested that
given a more general problem to solve after being trained to solve
similar problems very very well the poor old NN does worse that it
would have done if it had 0 training in the first place.
It's not like we dont know how to improve this kind of performance.
Try less hard in the training phase or make it "more noisy".
Turns out genetic methods are just the ticket for this.
The training produces less over-fitting and the resulting system
generalizes better than it did before training and more importantly
it takes maybe an order of magnitude crunching to produce a good answer
than the usual over-fit answer.
Anyway. Have to go and feed the cat.
Hi,
Come on Horsy Boy, you can do better. I
no where wrote something about curve
fitting and/or increasing the precision of
float point numbers:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
What makes you think LIPS measures precision?
You should know better as a 50% Prologer.
I explictily wrote here what the goal is:
"shave off some of the TOPS to do Prolog inferencing"
What are TOPS? Its a metric for GPUs:
TOPS stands for rCLTrillions of Operations Per Second.rCY https://www.lenovo.com/us/en/glossary/tops-in-computing/
See for yourself what is behind my post:
11.4 Giga Lips with a Budget Laptop
At the end of 2025 we acquired a couple of AI Laptops , that were still cheap, since RAM prices had not yet rocketed. The intend was to tap into
the Copilot+ certified hardware, and shave off some of the TOPS to do
Prolog inferencing. Amazingly our -C-WAM can churn 11.4 GIGA LIPS.
GPUs have evolved form lock-step to independent thread scheduling. This
made it possible to port the Hack VM variant, that forms the basis for
our -C-WAM, to WebGPU computer shaders. Using NUM_SHADERS = 4096 we could produce 11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M.
See also:
Medium Article - 11.4 Giga Lips
https://medium.com/2989/899b0d5c027b
So just get lost with your crazy irrelevant rant.
When I get more LIPS, things run faster, and
I remove digits from the time dimension.
Got it. Or are you too stupid?
Bye
R Kym Horsell schrieb:
In comp.lang.prolog Ross Finlayson <ross.a.finlayson@gmail.com> wrote:
...
Data centers should pay a 10000% excise on electricity,
wherever it comes from, a natural regulator of inverted economies.
And by ten thousand percent I really mean a ten thousand percent.
And what would a huge surcharge do?
Almost always end up affecting the less powerful end of society
with increased costs to services the AI industry will be doing
more and more of over time.
I started a little data center (exaflops.com) many years ago.
In those distant days people (in fact one was a prof of computer
science) told me you could never make money running a supercomputer.
LOL. :)
I've had many years to watch the trends and a far more efficient
way to solve resource problems in this area is to change the
algorithms. There is vast room for improvement, mostly because
of prevailing attitudes.
I used to do competetion data science as a sideline. Companies
would pay almost any price to get an extra decimal place in
the accuracy of their forecasting processes. But typically
they were trying to supercharge a system that should be scrapped
and re-designed from scratch. One area I'm thinking of is
investment. I had a customer one time -- like many times --
ask to improve a system that predicted the future price of
various stocks. The idea (for them) was to have as accurate a
prediction of what some stock would be worth in a week or a month's
time so that some moron could use the information to decide when
to buy or sell the thing.
I tried to argue the efficient thing was to create a system that
takes the human out of the loop altogether. It doesnt provide info
for someone to decide whether or not to follow the advice --
that is just introducing more noise into the loop and probably
cancels any benefit of adding a couple decimal places of precision.
What you *should* do is make a system that is tuned to robustly
maximize the profit from managing a portfolio.
Of course they wouldnt come at that. You can't suggest taking the
managers out of the loop. :)
Another idea relevant to current AI methods might be to curtail
use of typical neural net algorithms. Many of them try to squeeze
the best performance of some NN during the training-a phase in
the hope the resulting system will generalize well enough to be useful
on new data. But there's kind-of a law that the harder you train
some system to perform a task well, the less well they can subsuently
perform a more general version of the same thing. It's amusing when
you look at the graphs of NN being trained and then tested that
given a more general problem to solve after being trained to solve
similar problems very very well the poor old NN does worse that it
would have done if it had 0 training in the first place.
It's not like we dont know how to improve this kind of performance.
Try less hard in the training phase or make it "more noisy".
Turns out genetic methods are just the ticket for this.
The training produces less over-fitting and the resulting system
generalizes better than it did before training and more importantly
it takes maybe an order of magnitude crunching to produce a good answer
than the usual over-fit answer.
Anyway. Have to go and feed the cat.
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
How it started:
Captain: Throw the switch, Scotty!
Enterprise: Cloaking Device makes it invisible
Spock: Military secrets are the most fleeting of all.
Kirk Escapes the Romulans - The Enterprise Incident https://www.youtube.com/watch?v=AusAGjwlql8
How its going:
CEO Jensen Huang said the company has rCLlargely
concededrCY ChinarCOs artificial intelligence chip
market to Huawei, as U.S. export restrictions
continue to reshape the global AI semiconductor landscape.
Bye--- Synchronet 3.22a-Linux NewsLink 1.2
P.S.: What does China do?
HuaweirCOs semiconductor chief He Tingbo at the IEEE
ISCAS 2026 conference, Huawei's Tau Scaling Law is a newly
introduced semiconductor design framework that
shifts the industryrCOs optimization focus from
geometric scaling (shrinking physical transistor
sizes) to temporal scaling (compressing signal
propagation delay).
Nvidia Gave Up China - 4 Days Later THIS Happened https://www.youtube.com/watch?v=dLLw-qADKSU
Hi,
How it started:
Captain: Throw the switch, Scotty!
Enterprise: Cloaking Device makes it invisible
Spock: Military secrets are the most fleeting of all.
Kirk Escapes the Romulans - The Enterprise Incident https://www.youtube.com/watch?v=AusAGjwlql8
How its going:
CEO Jensen Huang said the company has rCLlargely
concededrCY ChinarCOs artificial intelligence chip
market to Huawei, as U.S. export restrictions
continue to reshape the global AI semiconductor landscape.
Bye--- Synchronet 3.22a-Linux NewsLink 1.2
P.S.: What does China do?
HuaweirCOs semiconductor chief He Tingbo at the IEEE
ISCAS 2026 conference, Huawei's Tau Scaling Law is a newly
introduced semiconductor design framework that
shifts the industryrCOs optimization focus from
geometric scaling (shrinking physical transistor
sizes) to temporal scaling (compressing signal
propagation delay).
Nvidia Gave Up China - 4 Days Later THIS Happened https://www.youtube.com/watch?v=dLLw-qADKSU
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Why implement both pre-emptive threadingAND cooperative tasks/engines?
Hi,
Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:
RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data
But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.
So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.
But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:
In-Browser LLM Inference Engine
https://webllm.mlc.ai/
My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.
But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese
hardware such some new Kirin AI in 2027.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Usual question:
Why implement both pre-emptive threadingAND cooperative tasks/engines?
I had implemented the ISO proposal in formerly Jekejeke
Prolog, you find the ISO proposal here:
ISO/IEC DTR 13211rCo5:2007
Prolog multi-threading support
https://logtalk.org/plstd/threads.pdf
But the ISO proposal doesn't match modern WebGPU APIs,
where your logical threads can live remotely in a dedicated GPU
in the VRAM there, and where you would have launch
parameters that say: Hey please run 4096 compute
shaders for me, that have independet thread state. Using
cooperative multi-tasking as the orchestrator works well.
Bye
Mild Shock schrieb:
Hi,
Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:
RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data
But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.
So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.
But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:
In-Browser LLM Inference Engine
https://webllm.mlc.ai/
My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.
But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese
hardware such some new Kirin AI in 2027.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Then the idea is that any of those can be found and matched in
one "run", i.e. a stall-less, branch-less, call-less list of less than
a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond.
So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,--- Synchronet 3.22a-Linux NewsLink 1.2
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched in
one "run", i.e. a stall-less, branch-less, call-less list of less than
a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond.
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
Hi,
Hurry Rossy Boy, the blue bus is waiting.
There is a quite a hyperbole from here:
Tesla S1070 in 2008
700 Watts , 1 Terra Flop
SOLVE TOMORROWrCOS PROBLEMS TODAY https://www.azken.com/download/Tesla_DS_S1070_EU.pdf
To here:
Blackwell GPU in 2026
575 Watts, 104.8 Terra Flops ( RTX 5090 )
From Volta To Blackwell https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell
But somehow the S1070 had already Massively-
Parallel, Many-Core Architecture, and forms
of MIMD, since it had 960 / 240 = 4 cores.
960 scalar processor cores (240 per GPU).
But possibly more resticted inside work
groups, than later NVIDIA Volta ff
architecture with independent thread state.
Bye
Disclaimer: The above is only a very rough
RTX 5090 spec. Its doesn't say what value
format and what vector/matrics ops were
used. Also energy consumption may vary.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched in
one "run", i.e. a stall-less, branch-less, call-less list of less than >> -a> a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond.
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
Hi,
So what does NUM_SHADERS = 4096 shaders mean here?
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Its only the number of logical threads.
CUDArao TEChNOLOGY UNLOCkS ThE POWER OF TESLA MANY-CORE PROCESSORS
The CUDA C compiler simplifies-a many-core programming
by enabling code development in a high-level language
and optimizing code to run on systems without knowledge of
how many cores are in the hardware.
CUDA applications automatically take advantage of more
cores or fewer cores in a system, so they can scale from
entry-level notebook GPUs to high end GPUs in technical
workstations-a and further into racks of GPUs in data
centers. This allows developers to
rCLcode oncerCY and deploy on a range of systems, as well as
scale forward in time as future GPUs deliver more
performance per watt and more cores per processor. The benefit
for software users is the opportunity to boost computing
performance simply by adding GPUs or using their
existing GPUs in new ways. https://www.azken.com/download/Tesla_DS_S1070_EU.pdf
Bye
Mild Shock schrieb:
Hi,
Hurry Rossy Boy, the blue bus is waiting.
There is a quite a hyperbole from here:
Tesla S1070 in 2008
700 Watts , 1 Terra Flop
SOLVE TOMORROWrCOS PROBLEMS TODAY
https://www.azken.com/download/Tesla_DS_S1070_EU.pdf
To here:
Blackwell GPU in 2026
575 Watts, 104.8 Terra Flops ( RTX 5090 )
-aFrom Volta To Blackwell
https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell
But somehow the S1070 had already Massively-
Parallel, Many-Core Architecture, and forms
of MIMD, since it had 960 / 240 = 4 cores.
960 scalar processor cores (240 per GPU).
But possibly more resticted inside work
groups, than later NVIDIA Volta ff
architecture with independent thread state.
Bye
Disclaimer: The above is only a very rough
RTX 5090 spec. Its doesn't say what value
format and what vector/matrics ops were
used. Also energy consumption may vary.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inthan
one "run", i.e. a stall-less, branch-less, call-less list of less
a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond.
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
.. bla bla goto bla bla ..
Stupid gangster: teamsters are a union.
In the trades, not the steals, ....
Hi,
If any of you guys do not understand what
is meant by or what the implications are:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Well I wouldn't care less. There are two
outcomes for numb nuts:
- Ignoramus: They don't understand it, but
-a they will understand it before they die.
- Ignorabimus: They don't understand it, and
-a will never understand it, and they die.
So who cares, its not my problem, you people
are stupid as fuck, and slow as fuck...
Bye
Mild Shock schrieb:
Hi,
Micro penis brain is in constant hiatus.
He can even not detect a trope.
LoL
Bye
Lane W schrieb:
Mild Shock wrote:miles away from her.
Hi,
My mother is worried that I fucked Lane W.
aka Micro Penis mother 24 hours straight.
She was screaming, basically singing all
the arias from operas that Luciano Pavarotti
usually sings. You Lane W. aka Micro Penis
should have heard it, since you
live in the basement of your mothers house.
No, actually remarkably, I don't. According to google I live 433
Strike!
See, what i said about you was spot on.
What you said about me was generic and incorrect.
You really suck, man.
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched in
one "run", i.e. a stall-less, branch-less, call-less list of less than
a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond.
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
I don't much care about Rust.
.. gibberish ..
Thief.
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched in
one "run", i.e. a stall-less, branch-less, call-less list of less than >> -a> a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond.
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inthan
one "run", i.e. a stall-less, branch-less, call-less list of less
a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond.
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,--- Synchronet 3.22a-Linux NewsLink 1.2
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Mild Shock wrote:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
post the link to the retailer, see what he sells for what money, idiot
The AMD Ryzen AI Halo PC is available for order exclusively through Micro Center in the United States. It is priced at $3,999.99
Hi,
This was archived on Jul 9, 2026:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Still, Jul 29, Rossy Boy halucinates accusations:
Ross Finlayson schrieb:
.. bla bla goto bla bla ..
Stupid gangster:-a teamsters are a union.
In the trades, not the steals, ....
Woa! Thats now 20 days of brain desease,
and not understanding the meaning and implications.
Even not understand pi-WAM has Hack VM backend.
But its all opensource. Bravo Rossy Boy, you are
champion in brainlessness and lazyness of
a idiot usenet troll.
Bye
Mild Shock schrieb:
Hi,
If any of you guys do not understand what
is meant by or what the implications are:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Well I wouldn't care less. There are two
outcomes for numb nuts:
- Ignoramus: They don't understand it, but
-a-a they will understand it before they die.
- Ignorabimus: They don't understand it, and
-a-a will never understand it, and they die.
So who cares, its not my problem, you people
are stupid as fuck, and slow as fuck...
Bye
Mild Shock schrieb:
Hi,
Micro penis brain is in constant hiatus.
He can even not detect a trope.
LoL
Bye
Lane W schrieb:
Mild Shock wrote:miles away from her.
Hi,
My mother is worried that I fucked Lane W.
aka Micro Penis mother 24 hours straight.
She was screaming, basically singing all
the arias from operas that Luciano Pavarotti
usually sings. You Lane W. aka Micro Penis
should have heard it, since you
live in the basement of your mothers house.
No, actually remarkably, I don't. According to google I live 433
Strike!
See, what i said about you was spot on.
What you said about me was generic and incorrect.
You really suck, man.
Hi,
This seems to be a funny Q16.16 experiment.
It shows that an integerish Hack can do
floatish stuff, by using binary fixpoint:
Raytracing on the Hack computer
2021/06/13 - im alex
https://blog.alexqua.ch/posts/from-nand-to-raytracer/
That it uses Rust is arbitrary. Feel free
to do it in C, C++, FORTRAN or Java. I guess
these languages all have basic arithmethic,
right? Maybe not a long jump always?
Bye
Mild Shock schrieb:
Hi,
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Just downloading some other person's code
Hi,
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inYou cannot make the mental translation that if you have:
one "run", i.e. a stall-less, branch-less, call-less list of less >>>> than
a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond. >>>>
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that >>>> -a> the indicators of the above as "positive presence" then is to make >>>> -a> for that the adjustments to the offsets and extents and the shifts >>>> -a> is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Just downloading some other person's code
I didn't do that, I wrote Hack VM for pi-WAM
from scratch, over the last 4 weeks. I came
back from holidays on end of June 2026, and now
we have end of July 2026. But its only possible
because the instruction set is very smal, like
ca. 8 functions and ca. 8 modes and ca. 8 conditions,
so its ca. 8 x 8 x 8 = 512 opcodes, each has an
A parameter and a D parameter simultaneously.
It has currently the following CPU backends:
-a- Now supports interleaved synchronous emulation.
-a- Now supports warp parallelism via Java platform threads.
-a- Now supports warp parallelism via Python system threads.
-a- Now supports warp parallelism via JavaScript worker threads.
-a- Note: For Python free threads are not yet fully tested.
-a- Note: For JavaScript web workers are not yet fully tested.
https://www.dogelog.ch/typtab/doclet/book/14_install/05_notes22/110_224.html
But frankly I came to encounter Hack not from
the usual university curriculum web resources,
but indirectly through a post about a Prolog
emulation of Hack, using constrained horn clauses (CHC):
Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/
The binary encoding is currently that the functions,
modes and conditions eat up a nibble (4-bit), in
total 12-bit, which I use then 10-bit for A parameter
and 10-bit for D parameter. I used AI freemium, Codex
by ChatGPT from within IntelliJ to do some fragment
code translations automatically from Java to JavaScript
or from JavaScript to Python.
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inless than
one "run", i.e. a stall-less, branch-less, call-less list of
a few or less than a few dozens or less than a few hundredsYou cannot make the mental translation that if you have:
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond. >>>>>
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that >>>>> -a> the indicators of the above as "positive presence" then is to make >>>>> -a> for that the adjustments to the offsets and extents and the shifts >>>>> -a> is according to those, otherwise no-ops. Then the idea is that a >>>>>As independent logical thread state, that automatically MIMD follows? >>>>>
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
The op-codes are all uniform, have the
same sub fields. Already Z-80 CPU differs here.
Another difference to a Z-80 CPU is that
their instruction stream was 8-bit, instructions
can 1, 2, 3 or 4 byte long. On the other
hand in my Hack VM all instructions are
one 32-bit chunk. The porting of a first
prototype that I already had, to WebGPU / WGSL
only took like 1-2 hours. The execution
of Hack VM is very simple, version 1.0,
for a single shader:
fn run() {
-a-a-a var pc : i32 = 0;
-a-a-a var accu : i32 = 0;
-a-a-a while (pc < i32(arrayLength(&code))) {
-a-a-a-a-a-a-a var instr : i32 = code[pc];
-a-a-a-a-a-a-a pc += 1;
-a-a-a-a-a-a-a var value : i32 = run_get(instr);
-a-a-a-a-a-a-a accu = run_fun(instr, accu, value);
-a-a-a-a-a-a-a run_set(instr, accu);
-a-a-a-a-a-a-a pc += run_jump(instr, accu);
-a-a-a }
}
https://github.com/Jean-Luc-Picard-2021/gigabudget/blob/b8946e891be774c40522267ab17062d32b023e7a/course/example63/boot.mjs#L176-L187
I first though this will be perfect for--- Synchronet 3.22a-Linux NewsLink 1.2
SIMD. Until I learnt that modern GPUs have
anyway MIMD. Hell Yeah, thats much better!
Bye
Mild Shock schrieb:
Hi,
Just downloading some other person's code
I didn't do that, I wrote Hack VM for pi-WAM
from scratch, over the last 4 weeks. I came
back from holidays on end of June 2026, and now
we have end of July 2026. But its only possible
because the instruction set is very smal, like
ca. 8 functions and ca. 8 modes and ca. 8 conditions,
so its ca. 8 x 8 x 8 = 512 opcodes, each has an
A parameter and a D parameter simultaneously.
It has currently the following CPU backends:
-a-a- Now supports interleaved synchronous emulation.
-a-a- Now supports warp parallelism via Java platform threads.
-a-a- Now supports warp parallelism via Python system threads.
-a-a- Now supports warp parallelism via JavaScript worker threads.
-a-a- Note: For Python free threads are not yet fully tested.
-a-a- Note: For JavaScript web workers are not yet fully tested.
https://www.dogelog.ch/typtab/doclet/book/14_install/05_notes22/110_224.html
But frankly I came to encounter Hack not from
the usual university curriculum web resources,
but indirectly through a post about a Prolog
emulation of Hack, using constrained horn clauses (CHC):
Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/
The binary encoding is currently that the functions,
modes and conditions eat up a nibble (4-bit), in
total 12-bit, which I use then 10-bit for A parameter
and 10-bit for D parameter. I used AI freemium, Codex
by ChatGPT from within IntelliJ to do some fragment
code translations automatically from Java to JavaScript
or from JavaScript to Python.
Have Fun!
Bye
Hi,
Just downloading some other person's code
I didn't do that, I wrote Hack VM for pi-WAM
from scratch, over the last 4 weeks. I came
back from holidays on end of June 2026, and now
we have end of July 2026. But its only possible
because the instruction set is very smal, like
ca. 8 functions and ca. 8 modes and ca. 8 conditions,
so its ca. 8 x 8 x 8 = 512 opcodes, each has an
A parameter and a D parameter simultaneously.
It has currently the following CPU backends:
-a- Now supports interleaved synchronous emulation.
-a- Now supports warp parallelism via Java platform threads.
-a- Now supports warp parallelism via Python system threads.
-a- Now supports warp parallelism via JavaScript worker threads.
-a- Note: For Python free threads are not yet fully tested.
-a- Note: For JavaScript web workers are not yet fully tested.
https://www.dogelog.ch/typtab/doclet/book/14_install/05_notes22/110_224.html
But frankly I came to encounter Hack not from
the usual university curriculum web resources,
but indirectly through a post about a Prolog
emulation of Hack, using constrained horn clauses (CHC):
Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/
The binary encoding is currently that the functions,
modes and conditions eat up a nibble (4-bit), in
total 12-bit, which I use then 10-bit for A parameter
and 10-bit for D parameter. I used AI freemium, Codex
by ChatGPT from within IntelliJ to do some fragment
code translations automatically from Java to JavaScript
or from JavaScript to Python.
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inless than
one "run", i.e. a stall-less, branch-less, call-less list of
a few or less than a few dozens or less than a few hundredsYou cannot make the mental translation that if you have:
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond. >>>>>
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that >>>>> -a> the indicators of the above as "positive presence" then is to make >>>>> -a> for that the adjustments to the offsets and extents and the shifts >>>>> -a> is according to those, otherwise no-ops. Then the idea is that a >>>>>As independent logical thread state, that automatically MIMD follows? >>>>>
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
There is Prolog compiler which spits out Hack.
From there on your are free to develop
and/or use any Hack realization that goes
from abstract to concrete. You could
replace the CPU backends that realize
a Hack VM by MIPS. Shouldn't be difficult.
Basically I refused to think in Huffman
Coding (*) while designing Hack VM. On the
other hand the MIPS architecture looks
like a big Huffman mess. Already its
initial design has 3 instructions types:
Type format (bits)
R opcode(6) rs(5) rt(5) rd(5) shamt(5) funct(6)
I opcode(6) rs(5) rt(5) imme(16)
J opcode(6) addr(26)
While my Hack has only 1 instruction
type, when binary encoded for Hack VM,
the currently used design looks as follows:
Type format (bits)
AD opcode(4) mode(4) cond(4) imme(10) addr(10)
But since its an abstract machine, nothing
prevents you from translating Hack code
into MIPS before executing it.
In has far you have to distinguish Hack,
which is specified in Prolog. And Hack VM
which is a virtual machine, with the above
instruction packing. And which has currently
a JavaScript runtime, a Python runtime
and a Java runtime.
Bye
(*)
https://en.wikipedia.org/wiki/Huffman_coding
Hi,
There is Prolog compiler which spits out Hack.
From there on your are free to develop
and/or use any Hack realization that goes
from abstract to concrete. You could
replace the CPU backends that realize
a Hack VM by MIPS. Shouldn't be difficult.
Basically I refused to think in Huffman
Coding (*) while designing Hack VM. On the
other hand the MIPS architecture looks
like a big Huffman mess. Already its
initial design has 3 instructions types:
Type format (bits)
R opcode(6) rs(5) rt(5) rd(5) shamt(5) funct(6)
I opcode(6) rs(5) rt(5) imme(16)
J opcode(6) addr(26)
While my Hack has only 1 instruction
type, when binary encoded for Hack VM,
the currently used design looks as follows:
Type format (bits)
AD opcode(4) mode(4) cond(4) imme(10) addr(10)
But since its an abstract machine, nothing
prevents you from translating Hack code
into MIPS before executing it.
In has far you have to distinguish Hack,
which is specified in Prolog. And Hack VM
which is a virtual machine, with the above
instruction packing. And which has currently
a JavaScript runtime, a Python runtime
and a Java runtime.
Bye
(*)
https://en.wikipedia.org/wiki/Huffman_coding
Hi,
There is Prolog compiler which spits out Hack.
From there on your are free to develop
and/or use any Hack realization that goes
from abstract to concrete. You could
replace the CPU backends that realize
a Hack VM by MIPS. Shouldn't be difficult.
Basically I refused to think in Huffman
Coding (*) while designing Hack VM. On the
other hand the MIPS architecture looks
like a big Huffman mess. Already its
initial design has 3 instructions types:
Type format (bits)
R opcode(6) rs(5) rt(5) rd(5) shamt(5) funct(6)
I opcode(6) rs(5) rt(5) imme(16)
J opcode(6) addr(26)
While my Hack has only 1 instruction
type, when binary encoded for Hack VM,
the currently used design looks as follows:
Type format (bits)
AD opcode(4) mode(4) cond(4) imme(10) addr(10)
But since its an abstract machine, nothing
prevents you from translating Hack code
into MIPS before executing it.
In has far you have to distinguish Hack,
which is specified in Prolog. And Hack VM
which is a virtual machine, with the above
instruction packing. And which has currently
a JavaScript runtime, a Python runtime
and a Java runtime.
Bye
(*)
https://en.wikipedia.org/wiki/Huffman_coding
Mild Shock schrieb:
Hi,
Just downloading some other person's code
I didn't do that, I wrote Hack VM for pi-WAM
from scratch, over the last 4 weeks. I came
back from holidays on end of June 2026, and now
we have end of July 2026. But its only possible
because the instruction set is very smal, like
ca. 8 functions and ca. 8 modes and ca. 8 conditions,
so its ca. 8 x 8 x 8 = 512 opcodes, each has an
A parameter and a D parameter simultaneously.
It has currently the following CPU backends:
- Now supports interleaved synchronous emulation.
- Now supports warp parallelism via Java platform threads.
- Now supports warp parallelism via Python system threads.
- Now supports warp parallelism via JavaScript worker threads.
- Note: For Python free threads are not yet fully tested.
- Note: For JavaScript web workers are not yet fully tested.
https://www.dogelog.ch/typtab/doclet/book/14_install/05_notes22/110_224.html >>
But frankly I came to encounter Hack not from
the usual university curriculum web resources,
but indirectly through a post about a Prolog
emulation of Hack, using constrained horn clauses (CHC):
Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/
The binary encoding is currently that the functions,
modes and conditions eat up a nibble (4-bit), in
total 12-bit, which I use then 10-bit for A parameter
and 10-bit for D parameter. I used AI freemium, Codex
by ChatGPT from within IntelliJ to do some fragment
code translations automatically from Java to JavaScript
or from JavaScript to Python.
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inless than
one "run", i.e. a stall-less, branch-less, call-less list of
a few or less than a few dozens or less than a few hundredsmicrosecond.
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents,that
the indicators of the above as "positive presence" then is to make >>>>>> > for that the adjustments to the offsets and extents and the shifts >>>>>> > is according to those, otherwise no-ops. Then the idea is that a >>>>>>As independent logical thread state, that automatically MIMD follows? >>>>>>
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
On 07/30/2026 01:32 PM, Mild Shock wrote:
Hi,
There is Prolog compiler which spits out Hack.
From there on your are free to develop
and/or use any Hack realization that goes
from abstract to concrete. You could
replace the CPU backends that realize
a Hack VM by MIPS. Shouldn't be difficult.
Basically I refused to think in Huffman
Coding (*) while designing Hack VM. On the
other hand the MIPS architecture looks
like a big Huffman mess. Already its
initial design has 3 instructions types:
Type format (bits)
R opcode(6) rs(5) rt(5) rd(5) shamt(5) funct(6)
I opcode(6) rs(5) rt(5) imme(16)
J opcode(6) addr(26)
While my Hack has only 1 instruction
type, when binary encoded for Hack VM,
the currently used design looks as follows:
Type format (bits)
AD opcode(4) mode(4) cond(4) imme(10) addr(10)
But since its an abstract machine, nothing
prevents you from translating Hack code
into MIPS before executing it.
In has far you have to distinguish Hack,
which is specified in Prolog. And Hack VM
which is a virtual machine, with the above
instruction packing. And which has currently
a JavaScript runtime, a Python runtime
and a Java runtime.
Bye
(*)
https://en.wikipedia.org/wiki/Huffman_coding
Mild Shock schrieb:
Hi,
Just downloading some other person's code
I didn't do that, I wrote Hack VM for pi-WAM
from scratch, over the last 4 weeks. I came
back from holidays on end of June 2026, and now
we have end of July 2026. But its only possible
because the instruction set is very smal, like
ca. 8 functions and ca. 8 modes and ca. 8 conditions,
so its ca. 8 x 8 x 8 = 512 opcodes, each has an
A parameter and a D parameter simultaneously.
It has currently the following CPU backends:
- Now supports interleaved synchronous emulation.
- Now supports warp parallelism via Java platform threads.
- Now supports warp parallelism via Python system threads.
- Now supports warp parallelism via JavaScript worker threads.
- Note: For Python free threads are not yet fully tested.
- Note: For JavaScript web workers are not yet fully tested.
https://www.dogelog.ch/typtab/doclet/book/14_install/05_notes22/110_224.html
But frankly I came to encounter Hack not from
the usual university curriculum web resources,
but indirectly through a post about a Prolog
emulation of Hack, using constrained horn clauses (CHC):
Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/
The binary encoding is currently that the functions,
modes and conditions eat up a nibble (4-bit), in
total 12-bit, which I use then 10-bit for A parameter
and 10-bit for D parameter. I used AI freemium, Codex
by ChatGPT from within IntelliJ to do some fragment
code translations automatically from Java to JavaScript
or from JavaScript to Python.
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched in >>>>>>> > one "run", i.e. a stall-less, branch-less, call-less list ofless than
a few or less than a few dozens or less than a few hundredsmicrosecond.
instructions, the results "findings" in data and corresponding >>>>>>> > "matchings" of expressions, that runs in less than one
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, >>>>>>> thatshifts
the indicators of the above as "positive presence" then is to >>>>>>> make
for that the adjustments to the offsets and extents and the
is according to those, otherwise no-ops. Then the idea is that a >>>>>>>As independent logical thread state, that automatically MIMD
follows?
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture >>>>>>>>>>> with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Huffman was a bright lad who it's said came up with Huffman
coding to get past taking a final arduous exam.
One of my grandfather's name is Huffman, he graduated high-school, somewhere's a picture of him with his high-school graduating
class, in Cleveland, Ohio, more than a hundred years ago.
Three people graduated high-school that year.
Huffman codes are a fantastic thing and ubiquitous, though,
like any other algorithm & data structure, they depend on
the distribution their best/average/worst cases, since everybody
here knows the general outline of asymptotics and Big O, Little o,
and Theta, about time and space and complexity.
Then, "adaptive algorithms" and "adapative data structures",
start with something like "compression" in usual accounts,
which usually is adaptive in the sense of compressing what's
compressible in a window, and making a histogram of the words
in the alphabet interpreted as a sequence of letters, and making
an optimal sort of Huffman-coding for entropy-coding for that,
since Huffman-coding is naturally optimal for a particular distribution,
that like power laws can be found in anything,
that one starts with making histograms after combinatorial enumeration,
and thus resulting an encoding into a language of bit-sequences
with the prefix-property that un-ambiguously makes for compressing
the compressible data.
Then the "adaptive" part of that is
periodically throwing that away and starting another.
One might aver that a more contextually-advised account
studies the entire corpus, for things like CCITT G4, JPEG,
JBIG, Deflate, and any account of entropy-coding or compression,
or Morse code, with of course both Huffman-coding and arithmetic-coding, since while Huffman-coding is obvious to everybody,
for a while some people thought arithmetic coding had patents,
which now are gone away, leaving all the above mentioned
and MPEG-4 also the "un-encumbered".
The notion of "summary statistics" and "order statistics" though,
in concrete mathematics of course is simple and clear.
Here's an article I read the other day about Huffman-coding,
I found it very insightful and quite enjoyable.
https://fgiesen.wordpress.com/2026/06/21/pivco-huffman-merge-operations/
Shut Up
On 07/31/2026 09:52 AM, Ross Finlayson wrote:
On 07/30/2026 01:32 PM, Mild Shock wrote:
Hi,
There is Prolog compiler which spits out Hack.
From there on your are free to develop
and/or use any Hack realization that goes
from abstract to concrete. You could
replace the CPU backends that realize
a Hack VM by MIPS. Shouldn't be difficult.
Basically I refused to think in Huffman
Coding (*) while designing Hack VM. On the
other hand the MIPS architecture looks
like a big Huffman mess. Already its
initial design has 3 instructions types:
Type format (bits)
R opcode(6) rs(5) rt(5) rd(5) shamt(5) funct(6)
I opcode(6) rs(5) rt(5) imme(16)
J opcode(6) addr(26)
While my Hack has only 1 instruction
type, when binary encoded for Hack VM,
the currently used design looks as follows:
Type format (bits)
AD opcode(4) mode(4) cond(4) imme(10) addr(10)
But since its an abstract machine, nothing
prevents you from translating Hack code
into MIPS before executing it.
In has far you have to distinguish Hack,
which is specified in Prolog. And Hack VM
which is a virtual machine, with the above
instruction packing. And which has currently
a JavaScript runtime, a Python runtime
and a Java runtime.
Bye
(*)
https://en.wikipedia.org/wiki/Huffman_coding
Mild Shock schrieb:
Hi,
Just downloading some other person's code
I didn't do that, I wrote Hack VM for pi-WAM
from scratch, over the last 4 weeks. I came
back from holidays on end of June 2026, and now
we have end of July 2026. But its only possible
because the instruction set is very smal, like
ca. 8 functions and ca. 8 modes and ca. 8 conditions,
so its ca. 8 x 8 x 8 = 512 opcodes, each has an
A parameter and a D parameter simultaneously.
It has currently the following CPU backends:
- Now supports interleaved synchronous emulation.
- Now supports warp parallelism via Java platform threads.
- Now supports warp parallelism via Python system threads.
- Now supports warp parallelism via JavaScript worker threads.
- Note: For Python free threads are not yet fully tested.
- Note: For JavaScript web workers are not yet fully tested.
https://www.dogelog.ch/typtab/doclet/book/14_install/05_notes22/110_224.html
But frankly I came to encounter Hack not from
the usual university curriculum web resources,
but indirectly through a post about a Prolog
emulation of Hack, using constrained horn clauses (CHC):
Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/
The binary encoding is currently that the functions,
modes and conditions eat up a nibble (4-bit), in
total 12-bit, which I use then 10-bit for A parameter
and 10-bit for D parameter. I used AI freemium, Codex
by ChatGPT from within IntelliJ to do some fragment
code translations automatically from Java to JavaScript
or from JavaScript to Python.
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms. >>>>>>> >
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched in >>>>>>>> > one "run", i.e. a stall-less, branch-less, call-less list of >>>>>>>> less thanmicrosecond.
a few or less than a few dozens or less than a few hundreds >>>>>>>> > instructions, the results "findings" in data and corresponding >>>>>>>> > "matchings" of expressions, that runs in less than one
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, >>>>>>>> thatAs independent logical thread state, that automatically MIMD
the indicators of the above as "positive presence" then is to >>>>>>>> make
for that the adjustments to the offsets and extents and the >>>>>>>> shifts
is according to those, otherwise no-ops. Then the idea is that a >>>>>>>>
follows?
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture >>>>>>>>>>>> with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Huffman was a bright lad who it's said came up with Huffman
coding to get past taking a final arduous exam.
One of my grandfather's name is Huffman, he graduated high-school,
somewhere's a picture of him with his high-school graduating
class, in Cleveland, Ohio, more than a hundred years ago.
Three people graduated high-school that year.
Huffman codes are a fantastic thing and ubiquitous, though,
like any other algorithm & data structure, they depend on
the distribution their best/average/worst cases, since everybody
here knows the general outline of asymptotics and Big O, Little o,
and Theta, about time and space and complexity.
Then, "adaptive algorithms" and "adapative data structures",
start with something like "compression" in usual accounts,
which usually is adaptive in the sense of compressing what's
compressible in a window, and making a histogram of the words
in the alphabet interpreted as a sequence of letters, and making
an optimal sort of Huffman-coding for entropy-coding for that,
since Huffman-coding is naturally optimal for a particular distribution,
that like power laws can be found in anything,
that one starts with making histograms after combinatorial enumeration,
and thus resulting an encoding into a language of bit-sequences
with the prefix-property that un-ambiguously makes for compressing
the compressible data.
Then the "adaptive" part of that is
periodically throwing that away and starting another.
One might aver that a more contextually-advised account
studies the entire corpus, for things like CCITT G4, JPEG,
JBIG, Deflate, and any account of entropy-coding or compression,
or Morse code, with of course both Huffman-coding and arithmetic-coding,
since while Huffman-coding is obvious to everybody,
for a while some people thought arithmetic coding had patents,
which now are gone away, leaving all the above mentioned
and MPEG-4 also the "un-encumbered".
The notion of "summary statistics" and "order statistics" though,
in concrete mathematics of course is simple and clear.
Here's an article I read the other day about Huffman-coding,
I found it very insightful and quite enjoyable.
https://fgiesen.wordpress.com/2026/06/21/pivco-huffman-merge-operations/
Shut Up
"It easily scales both up and down
with the capabilities of the target machine."
-
https://fgiesen.wordpress.com/2026/06/21/pivco-huffman-merge-operations/
Hi,
Just downloading some other person's code
I didn't do that, I wrote Hack VM for pi-WAM
from scratch, over the last 4 weeks. I came
back from holidays on end of June 2026, and now
we have end of July 2026. But its only possible
because the instruction set is very smal, like
ca. 8 functions and ca. 8 modes and ca. 8 conditions,
so its ca. 8 x 8 x 8 = 512 opcodes, each has an
A parameter and a D parameter simultaneously.
It has currently the following CPU backends:
-a- Now supports interleaved synchronous emulation.
-a- Now supports warp parallelism via Java platform threads.
-a- Now supports warp parallelism via Python system threads.
-a- Now supports warp parallelism via JavaScript worker threads.
-a- Note: For Python free threads are not yet fully tested.
-a- Note: For JavaScript web workers are not yet fully tested.
https://www.dogelog.ch/typtab/doclet/book/14_install/05_notes22/110_224.html
But frankly I came to encounter Hack not from
the usual university curriculum web resources,
but indirectly through a post about a Prolog
emulation of Hack, using constrained horn clauses (CHC):
Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/
The binary encoding is currently that the functions,
modes and conditions eat up a nibble (4-bit), in
total 12-bit, which I use then 10-bit for A parameter
and 10-bit for D parameter. I used AI freemium, Codex
by ChatGPT from within IntelliJ to do some fragment
code translations automatically from Java to JavaScript
or from JavaScript to Python.
Have Fun!
Bye
Mild Shock schrieb:
Hi,
Who exactly is the thief? Does this person
have stats in the Rogue class in dungeons
and dragons?
The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.
Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,
and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.
The gist are projects like here, first to
educate yourself about Hack:
https://www.nand2tetris.org/course
And then to use Hack in different contexts:
https://www.nand2tetris.org/copy-of-talks
For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,
has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.
Bye
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inless than
one "run", i.e. a stall-less, branch-less, call-less list of
a few or less than a few dozens or less than a few hundredsYou cannot make the mental translation that if you have:
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond. >>>>>
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that >>>>> -a> the indicators of the above as "positive presence" then is to make >>>>> -a> for that the adjustments to the offsets and extents and the shifts >>>>> -a> is according to those, otherwise no-ops. Then the idea is that a >>>>>As independent logical thread state, that automatically MIMD follows? >>>>>
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Mostlikely for high performance computing |a la,
the Actor/Erlang model is dead, they might rely
on MPMC (Multiple Producer, Multiple Consumer)
queue entities separate from the threads. The
ISO Prolog multi-threading support had also such
threads. But besides that was also Actor/Erlang
leaning in practice, like SWI, where threads
have some default queues. So an actor is basically
a Thread and Mailbox conflation. While a MPMC queue
is a kind of separate Mailbox, where multiple
"actors" can read from and write from. A kind of
localized Linda Tuple store.
Which Programming language did adopted the
non-Actor pi-calculus model? Right golang
with its channels.
Bye
Mild Shock schrieb:
Hi,
Usual question:
Why implement both pre-emptive threadingAND cooperative tasks/engines?
I had implemented the ISO proposal in formerly Jekejeke
Prolog, you find the ISO proposal here:
ISO/IEC DTR 13211rCo5:2007
Prolog multi-threading support
https://logtalk.org/plstd/threads.pdf
But the ISO proposal doesn't match modern WebGPU APIs,
where your logical threads can live remotely in a dedicated GPU
in the VRAM there, and where you would have launch
parameters that say: Hey please run 4096 compute
shaders for me, that have independet thread state. Using
cooperative multi-tasking as the orchestrator works well.
Bye
Mild Shock schrieb:
Hi,
Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:
RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data
But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.
So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.
But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:
In-Browser LLM Inference Engine
https://webllm.mlc.ai/
My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.
But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese
hardware such some new Kirin AI in 2027.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:
RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data
But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.
So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.
But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:
In-Browser LLM Inference Engine
https://webllm.mlc.ai/
My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.
But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese
hardware such some new Kirin AI in 2027.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Tablets and phone are more annoying to
use with WebGPU. The usual browsers don't
have a Chrome DevTools panel integrated,
so that one could do JavaScript Debugging
directly on the device. Instead one has to
use a desktop machine, and connect the
device via UBS-C , and start a Chrome
Browser there . And then start a Chrome
DevTools panel alone, that is pair with
the device, via UBS-C cable. So this way
I already see where it crashes on the
tablets and phone:
await output.mapAsync(GPUMapMode.READ)
Unhandled Promise Rejection: OperationError
The above is the error that one can re-produce
already here with this test:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Not sure what exactly happens. Maybe
a form of timeout or device lost, that the
primitive HTML / JavaScript doesn't handle
gracefully yet. Maybe redimensioning the
test, so that it consumes less time would
help. Who knows? Will see. For production
use of a GPU integration I have to anyway
provide work slicing it seems.
Bye
Mild Shock schrieb:
Hi,
Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:
RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data
But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.
So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.
But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:
In-Browser LLM Inference Engine
https://webllm.mlc.ai/
My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.
But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese
hardware such some new Kirin AI in 2027.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,--- Synchronet 3.22a-Linux NewsLink 1.2
Looking at the floor plan of a NPU:
Getting peak TOPS on a Ryzen AI 7 350 NPU https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/
It seems to me comms between tiles takes
at least Manhattan Distance or L1 Norm time,
if there is no comms congestion
But how does a packet travel? This way:
+----E
|
|
S
Or this way, from start S to end E:
-a-a +-E
-a +
-a+
S
And what does the chip do if there is
traffic congestion? Some papers are
here, possibly an old problem giving
that processor "cubes" are nothing new.
But a "cube" would be 3D and not 2D.
This paper is old from 2007 or so:
Routing Algorithms for 2D NoC Architectures http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
Hi,
Tablets and phone are more annoying to
use with WebGPU. The usual browsers don't
have a Chrome DevTools panel integrated,
so that one could do JavaScript Debugging
directly on the device. Instead one has to
use a desktop machine, and connect the
device via UBS-C , and start a Chrome
Browser there . And then start a Chrome
DevTools panel alone, that is pair with
the device, via UBS-C cable. So this way
I already see where it crashes on the
tablets and phone:
await output.mapAsync(GPUMapMode.READ)
Unhandled Promise Rejection: OperationError
The above is the error that one can re-produce
already here with this test:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Not sure what exactly happens. Maybe
a form of timeout or device lost, that the
primitive HTML / JavaScript doesn't handle
gracefully yet. Maybe redimensioning the
test, so that it consumes less time would
help. Who knows? Will see. For production
use of a GPU integration I have to anyway
provide work slicing it seems.
Bye
Of course we can make a special texture to handle it.
On 8/1/2026 5:22 AM, Mild Shock wrote:shader. Of course we can make a special texture to handle it. But, we
Hi,[...]
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You are going to need a place to allocate nodes in the compute
Hi,
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You have not only start
S to end E communication:
+----E
|
|
S
You might also have ACK or NACK
from E or midpoints back to S:
-a-a S'
-a +
-a+
E'
Ok, I made that up, I have no idea what a flit is,
when the author wrote this here:
"Packet flits are held in the FIFO which can
be used to determine back pressure. Dropping flits
in a NoC may not be possible since these
architectures may not provide an end-to-end
protocol for retransmission."
Routing Algorithms for 2D NoC Architectures http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
Mild Shock schrieb:
Hi,
Looking at the floor plan of a NPU:
Getting peak TOPS on a Ryzen AI 7 350 NPU
https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/
It seems to me comms between tiles takes
at least Manhattan Distance or L1 Norm time,
if there is no comms congestion
But how does a packet travel? This way:
+----E
|
|
S
Or this way, from start S to end E:
-a-a-a +-E
-a-a +
-a-a+
S
And what does the chip do if there is
traffic congestion? Some papers are
here, possibly an old problem giving
that processor "cubes" are nothing new.
But a "cube" would be 3D and not 2D.
This paper is old from 2007 or so:
Routing Algorithms for 2D NoC Architectures
http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
It has textures to work with in the pipeline.
Hi,
WebGPU and WebGL are two different things. I explained
that towards you already like 3-5 times.
Of course we can make a special texture to handle it.
You still don't understand that I am using WebGPU,
and not WebGL. WebGPU has three improvements,
that from your talking are missing in WebGL?
- It has compute shaders
- It has arrays
- It has structs
- What else?
I didn't use structs in my example, although Gemini
nearly forced me to use structs. But you could
use a struct with fields and some of these arrays
to represent a queue. But here in this example
that is open source, I only used flat arrays. I
nowhere needed to abuse textures to store something:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
You can study the source code, the arrays have
CUDA inspired binding annotations but are not
CUDA but rather WGSL:
Hack VM as a Compute Shader in WGSL
@group(0) @binding(0) var<storage, read> code: array<i32>;
@group(0) @binding(1) var<storage, read_write> state: array<i32>; https://github.com/Jean-Luc-Picard-2021/gigabudget/blob/main/course/example63/boot.mjs
You can say whether a buffer is read, or read_write.
Buffers can be transfered from CPU to GPU, before
running commands, and transfered back from GPU to
CPU after running commands. The use case that
you find on GitHub uses both. Namely also fetching
results via a buffer, to then show them in
the HTML page. Shouldn't be much a problem to
run the example at home locally, all you need is
a HTTPS server. But the example is not yet queues,
but it already shows the foundation, which is WebGPU
with its language WGSL and not WebGL with its language
GLSL. These are two different things.
I explained that towards you already like 3-5 times.
Bye
Chris M. Thomasson schrieb:
On 8/1/2026 5:22 AM, Mild Shock wrote:
Hi,[...]
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You are going to need a place to allocate nodes in the computeshader. Of course we can make a special texture to handle it. But, we
need to strive to avoid a wait condition. I don't want a compute shader
to spin. Yes, CAS can be used, but, try to make it be used as a "state machine", where the transitions from states are atomic. Try to avoid it making a loop, where we loop on failure.
Mild Shock schrieb:
Hi,
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You have not only start
S to end E communication:
+----E
|
|
S
You might also have ACK or NACK
from E or midpoints back to S:
-a-a-a S'
-a-a +
-a-a+
E'
Ok, I made that up, I have no idea what a flit is,
when the author wrote this here:
"Packet flits are held in the FIFO which can
be used to determine back pressure. Dropping flits
in a NoC may not be possible since these
architectures may not provide an end-to-end
protocol for retransmission."
Routing Algorithms for 2D NoC Architectures
http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
Mild Shock schrieb:
Hi,
Looking at the floor plan of a NPU:
Getting peak TOPS on a Ryzen AI 7 350 NPU
https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/
It seems to me comms between tiles takes
at least Manhattan Distance or L1 Norm time,
if there is no comms congestion
But how does a packet travel? This way:
+----E
|
|
S
Or this way, from start S to end E:
-a-a-a +-E
-a-a +
-a-a+
S
And what does the chip do if there is
traffic congestion? Some papers are
here, possibly an old problem giving
that processor "cubes" are nothing new.
But a "cube" would be 3D and not 2D.
This paper is old from 2007 or so:
Routing Algorithms for 2D NoC Architectures
http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
But, I still don't know what you main goal is?
shave off some of the TOPS to do Prolog inferencing
Hi,
It has textures to work with in the pipeline.
Hi,
Why would I use text inside my compute shader.
Could you tell me. The Hack VM doesn't do
textures. You are confused. There is nothing
about textures here:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
You can read the text , it says nowhere
consume or produce textures. Its not a rendering
application. I use the compute shader to run Prolog:
"At the end of 2025 we acquired a couple of
AI Laptops , that were still cheap, since
RAM prices had not yet rocketed. The intend
was to tap into the Copilot+ certified hardware,
and shave off some of the TOPS to do Prolog
inferencing. Amazingly our -C-WAM can
churn 11.4 GIGA LIPS.
GPUs have evolved form lock-step to independent
thread scheduling. This made it possible to
port the Hack VM variant, that forms the basis
for our -C-WAM, to WebGPU computer shaders.
Using NUM_SHADERS = 4096 we could produce
11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M."
Bye
Mild Shock schrieb:
Hi,
WebGPU and WebGL are two different things. I explained
that towards you already like 3-5 times.
Of course we can make a special texture to handle it.
You still don't understand that I am using WebGPU,
and not WebGL. WebGPU has three improvements,
that from your talking are missing in WebGL?
- It has compute shaders
- It has arrays
- It has structs
- What else?
I didn't use structs in my example, although Gemini
nearly forced me to use structs. But you could
use a struct with fields and some of these arrays
to represent a queue. But here in this example
that is open source, I only used flat arrays. I
nowhere needed to abuse textures to store something:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
You can study the source code, the arrays have
CUDA inspired binding annotations but are not
CUDA but rather WGSL:
Hack VM as a Compute Shader in WGSL
@group(0) @binding(0) var<storage, read> code: array<i32>;
@group(0) @binding(1) var<storage, read_write> state: array<i32>;
https://github.com/Jean-Luc-Picard-2021/gigabudget/blob/main/course/example63/boot.mjs
You can say whether a buffer is read, or read_write.
Buffers can be transfered from CPU to GPU, before
running commands, and transfered back from GPU to
CPU after running commands. The use case that
you find on GitHub uses both. Namely also fetching
results via a buffer, to then show them in
the HTML page. Shouldn't be much a problem to
run the example at home locally, all you need is
a HTTPS server. But the example is not yet queues,
but it already shows the foundation, which is WebGPU
with its language WGSL and not WebGL with its language
GLSL. These are two different things.
I explained that towards you already like 3-5 times.
Bye
Chris M. Thomasson schrieb:
On 8/1/2026 5:22 AM, Mild Shock wrote:shader. Of course we can make a special texture to handle it. But, we
Hi,[...]
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You are going to need a place to allocate nodes in the compute
need to strive to avoid a wait condition. I don't want a compute
shader to spin. Yes, CAS can be used, but, try to make it be used as a
"state machine", where the transitions from states are atomic. Try to
avoid it making a loop, where we loop on failure.
Mild Shock schrieb:
Hi,
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You have not only start
S to end E communication:
+----E
|
|
S
You might also have ACK or NACK
from E or midpoints back to S:
-a-a-a S'
-a-a +
-a-a+
E'
Ok, I made that up, I have no idea what a flit is,
when the author wrote this here:
"Packet flits are held in the FIFO which can
be used to determine back pressure. Dropping flits
in a NoC may not be possible since these
architectures may not provide an end-to-end
protocol for retransmission."
Routing Algorithms for 2D NoC Architectures
http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
Mild Shock schrieb:
Hi,
Looking at the floor plan of a NPU:
Getting peak TOPS on a Ryzen AI 7 350 NPU
https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/ >>>>
It seems to me comms between tiles takes
at least Manhattan Distance or L1 Norm time,
if there is no comms congestion
But how does a packet travel? This way:
+----E
|
|
S
Or this way, from start S to end E:
-a-a-a +-E
-a-a +
-a-a+
S
And what does the chip do if there is
traffic congestion? Some papers are
here, possibly an old problem giving
that processor "cubes" are nothing new.
But a "cube" would be 3D and not 2D.
This paper is old from 2007 or so:
Routing Algorithms for 2D NoC Architectures
http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
But, I still don't know what you main goal is?The goal is "Prolog inferencing"
It has textures to work with in the pipeline.I don't need textures for "Prolog inferencing"
On 8/1/26 05:19, Mild Shock wrote:
Hi,
Tablets and phone are more annoying to
use with WebGPU. The usual browsers don't
have a Chrome DevTools panel integrated,
so that one could do JavaScript Debugging
directly on the device. Instead one has to
use a desktop machine, and connect the
device via UBS-C , and start a Chrome
Browser there . And then start a Chrome
DevTools panel alone, that is pair with
the device, via UBS-C cable. So this way
I already see where it crashes on the
tablets and phone:
await output.mapAsync(GPUMapMode.READ)
Unhandled Promise Rejection: OperationError
The above is the error that one can re-produce
already here with this test:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Not sure what exactly happens. Maybe
a form of timeout or device lost, that the
primitive HTML / JavaScript doesn't handle
gracefully yet. Maybe redimensioning the
test, so that it consumes less time would
help. Who knows? Will see. For production
use of a GPU integration I have to anyway
provide work slicing it seems.
Bye
I know that you are Hanson and I know that you are a cocksucker.
My question is, though, do you know, do, vomit, defecate, smell,
evaporate, sweat, ooze in, ooze out, fart, see, hear, sense, taste, and GESTATE anything other than programming?
Programming is just a tool, you know. It is nothing by itself worth even mentioning.
node.exe dogelog.mjsDogelog Spieler 2.2.5, Node, JavaScript 26.4.0
Hi,
Chris M. Thomasson can ask 100 more questions.
I will happily answer them. But maybe I should
make a Wiki to explain the ever same things:
But, I still don't know what you main goal is?The goal is "Prolog inferencing"
It has textures to work with in the pipeline.I don't need textures for "Prolog inferencing"
98 more questions to go, don't give up!
Bye
Taskfreak schrieb:
On 8/1/26 05:19, Mild Shock wrote:
Hi,
Tablets and phone are more annoying to
use with WebGPU. The usual browsers don't
have a Chrome DevTools panel integrated,
so that one could do JavaScript Debugging
directly on the device. Instead one has to
use a desktop machine, and connect the
device via UBS-C , and start a Chrome
Browser there . And then start a Chrome
DevTools panel alone, that is pair with
the device, via UBS-C cable. So this way
I already see where it crashes on the
tablets and phone:
await output.mapAsync(GPUMapMode.READ)
Unhandled Promise Rejection: OperationError
The above is the error that one can re-produce
already here with this test:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Not sure what exactly happens. Maybe
a form of timeout or device lost, that the
primitive HTML / JavaScript doesn't handle
gracefully yet. Maybe redimensioning the
test, so that it consumes less time would
help. Who knows? Will see. For production
use of a GPU integration I have to anyway
provide work slicing it seems.
Bye
I know that you are Hanson and I know that you are a cocksucker.
My question is, though, do you know, do, vomit, defecate, smell,
evaporate, sweat, ooze in, ooze out, fart, see, hear, sense, taste,
and GESTATE anything other than programming?
Programming is just a tool, you know. It is nothing by itself worth
even mentioning.
On 8/1/2026 5:47 PM, Mild Shock wrote:building vector fields, etc.... And yes I use textures for some input
Hi,[...]
Chris M. Thomasson can ask 100 more questions.
I will happily answer them. But maybe I should
make a Wiki to explain the ever same things:
But, I still don't know what you main goal is?The goal is "Prolog inferencing"
It has textures to work with in the pipeline.I don't need textures for "Prolog inferencing"
98 more questions to go, don't give up!
Fwiw, I have several compute shaders that do what I want. Mainly
For instance, this is 100% wait free.
void add_hit(ct_plane2d plane, vec2 p, vec3 weight)
{
vec2 uv = ct_plane2d_unproject(plane, p);
ivec2 px = ivec2(uv * u_resolution);
if (px.x >= 0 && px.x < int(u_resolution.x) &&
px.y >= 0 && px.y < int(u_resolution.y))
{
imageAtomicAdd(accum_r, px, weight.r);
imageAtomicAdd(accum_g, px, weight.g);
imageAtomicAdd(accum_b, px, weight.b);
imageAtomicAdd(accum_hits, px, 1.0f);
}
}
Notice how I separated my accumulation buffer into different textures?alpha / hit counter
layout(binding = 0, r32f) uniform coherent image2D accum_r;
layout(binding = 1, r32f) uniform coherent image2D accum_g;
layout(binding = 2, r32f) uniform coherent image2D accum_b;
layout(binding = 3, r32f) uniform coherent image2D accum_hits; //
Works great and runs really fast.
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Strive to never make a compute shader wait
on something, like an empty condition of a queue, stack.
Hi,
Ok, following the instructions here:
npm install webgpu
https://github.com/dawn-gpu/node-webgpu
I can now run webgpu also from CLI:
node.exe dogelog.mjsDogelog Spieler 2.2.5, Node, JavaScript 26.4.0
(c) 1985-2026, XLOG Technologies AG, Schweiz
?- ensure_loaded(library(edge/furryhaze)).
true.
?- between(1,3,_), time(expedite((between(1,100,_),
between(1,100,_), between(1,100,_)), [size(4096)])), fail.
% Zeit 1037.994 ms, GC 0.000 ms, Lips 111 k
% Zeit 1091.131 ms, GC 0.000 ms, Lips 106 k
% Zeit 1045.274 ms, GC 0.000 ms, Lips 110 k
fail.
Same benchmark result as in the browser.
Now I can rent a bigger GPU by the hour
and do some easy CLI testing.
LoL
Bye
Mild Shock schrieb:
Hi,
Chris M. Thomasson can ask 100 more questions.
I will happily answer them. But maybe I should
make a Wiki to explain the ever same things:
But, I still don't know what you main goal is?The goal is "Prolog inferencing"
It has textures to work with in the pipeline.I don't need textures for "Prolog inferencing"
98 more questions to go, don't give up!
Bye
Taskfreak schrieb:
On 8/1/26 05:19, Mild Shock wrote:
Hi,
Tablets and phone are more annoying to
use with WebGPU. The usual browsers don't
have a Chrome DevTools panel integrated,
so that one could do JavaScript Debugging
directly on the device. Instead one has to
use a desktop machine, and connect the
device via UBS-C , and start a Chrome
Browser there . And then start a Chrome
DevTools panel alone, that is pair with
the device, via UBS-C cable. So this way
I already see where it crashes on the
tablets and phone:
await output.mapAsync(GPUMapMode.READ)
Unhandled Promise Rejection: OperationError
The above is the error that one can re-produce
already here with this test:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Not sure what exactly happens. Maybe
a form of timeout or device lost, that the
primitive HTML / JavaScript doesn't handle
gracefully yet. Maybe redimensioning the
test, so that it consumes less time would
help. Who knows? Will see. For production
use of a GPU integration I have to anyway
provide work slicing it seems.
Bye
I know that you are Hanson and I know that you are a cocksucker.
My question is, though, do you know, do, vomit, defecate, smell,
evaporate, sweat, ooze in, ooze out, fart, see, hear, sense, taste,
and GESTATE anything other than programming?
Programming is just a tool, you know. It is nothing by itself worth
even mentioning.
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Hi,
This was archived on Jul 9, 2026:
11.4 Giga Lips with a Budget Laptop https://github.com/Jean-Luc-Picard-2021/gigabudget
Still, Jul 29, Rossy Boy halucinates accusations:
Ross Finlayson schrieb:
.. bla bla goto bla bla ..
Stupid gangster:-a teamsters are a union.
In the trades, not the steals, ....
Woa! Thats now 20 days of brain desease,
and not understanding the meaning and implications.
Even not understand pi-WAM has Hack VM backend.
But its all opensource. Bravo Rossy Boy, you are
champion in brainlessness and lazyness of
a idiot usenet troll.
Bye
Mild Shock schrieb:
Hi,
If any of you guys do not understand what
is meant by or what the implications are:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
Well I wouldn't care less. There are two
outcomes for numb nuts:
- Ignoramus: They don't understand it, but
-a-a they will understand it before they die.
- Ignorabimus: They don't understand it, and
-a-a will never understand it, and they die.
So who cares, its not my problem, you people
are stupid as fuck, and slow as fuck...
Bye
Mild Shock schrieb:
Hi,
Micro penis brain is in constant hiatus.
He can even not detect a trope.
LoL
Bye
Lane W schrieb:
Mild Shock wrote:miles away from her.
Hi,
My mother is worried that I fucked Lane W.
aka Micro Penis mother 24 hours straight.
She was screaming, basically singing all
the arias from operas that Luciano Pavarotti
usually sings. You Lane W. aka Micro Penis
should have heard it, since you
live in the basement of your mothers house.
No, actually remarkably, I don't. According to google I live 433
Strike!
See, what i said about you was spot on.
What you said about me was generic and incorrect.
You really suck, man.
Hi,
Do a YouTube video about it:
Topic: Tit for Tat, or how I messed up
with an innocent poster, and learnt about FAFO:
#fuckaroundandfindout
https://www.youtube.com/shorts/6ALRRksc72M
You were provable the first idiot, posting
stupid comments into my posts, besides of
course Micro Penis, who is a paid troll.
Have Fun!
Bye
Ross Finlayson schrieb:
For dummies, ....
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inthan
one "run", i.e. a stall-less, branch-less, call-less list of less
a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond.
You cannot make the mental translation that if you have:
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
of various approaches to Szemeredi, and about the independence
of various approaches of entropy, or Aristotle and Leibniz
wrote Newton's method, where of course Kepler wrote
the System of the World's universal gravitation, that
Hi,
Or do a YouTube video about:
Standing on the shoulders of giants https://en.wikipedia.org/wiki/Standing_on_the_shoulders_of_giants
Calling people who build software "thieves",
is probably the most philosopher syphilis brain
thing I ever heard in 2026. You should really
jump from a bridge Rossy Boy. I think its over
for you, the lamps have already gone out...
Bye
Mild Shock schrieb:
Hi,
Do a YouTube video about it:
Topic: Tit for Tat, or how I messed up
with an innocent poster, and learnt about FAFO:
#fuckaroundandfindout
https://www.youtube.com/shorts/6ALRRksc72M
You were provable the first idiot, posting
stupid comments into my posts, besides of
course Micro Penis, who is a paid troll.
Have Fun!
Bye
Ross Finlayson schrieb:
For dummies, ....
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inYou cannot make the mental translation that if you have:
one "run", i.e. a stall-less, branch-less, call-less list of less >>>> than
a few or less than a few dozens or less than a few hundreds
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond. >>>>
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that >>>> -a> the indicators of the above as "positive presence" then is to make >>>> -a> for that the adjustments to the offsets and extents and the shifts >>>> -a> is according to those, otherwise no-ops. Then the idea is that a
As independent logical thread state, that automatically MIMD follows?
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
RF rCo transcript received and read. The session closed well. What we
mapped out across these rounds is, to my mind, a credible foundation:
two matcher normal forms (AND for properties, XOR for code-points), a validated SSE2 smearing sequence via Claude's S1/S2 sketch, a closed
calling convention with explicit ABI spill gates, and the SBC-less
design mantra as a gradient rather than a boolean. The open items rCo
stack tagging, AST wire format, bit-granular Viswath boundaries rCo are properly scoped for next time rather than lost.
Hi,
Ross Finlayson schrieb:
of various approaches to Szemeredi, and about the independence
of various approaches of entropy, or Aristotle and Leibniz
You horrible horrible Person and Thief.
Balantly stealing from Szemeredi, Aristotle,
Leibniz, etc..
Ross Finlayson schrieb:
wrote Newton's method, where of course Kepler wrote
the System of the World's universal gravitation, that
And poor Newton and Kepler get also exploited,
from shameless Rossy Boy. Thats not very original,
shame on you!
I guess this is the final verdict for you. As a
person without original thought, you need
to do as a big favor,
and jump from a bridge.
Bye
Mild Shock schrieb:
Hi,
Or do a YouTube video about:
Standing on the shoulders of giants
https://en.wikipedia.org/wiki/Standing_on_the_shoulders_of_giants
Calling people who build software "thieves",
is probably the most philosopher syphilis brain
thing I ever heard in 2026. You should really
jump from a bridge Rossy Boy. I think its over
for you, the lamps have already gone out...
Bye
Mild Shock schrieb:
Hi,
Do a YouTube video about it:
Topic: Tit for Tat, or how I messed up
with an innocent poster, and learnt about FAFO:
#fuckaroundandfindout
https://www.youtube.com/shorts/6ALRRksc72M
You were provable the first idiot, posting
stupid comments into my posts, besides of
course Micro Penis, who is a paid troll.
Have Fun!
Bye
Ross Finlayson schrieb:
For dummies, ....
Mild Shock schrieb:
Hi,
I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,
enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.
Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you
with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for
you alone, your stupid asshole.
Bye
Ross Finlayson schrieb:
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
I don't much care about Rust.
.. gibberish ..
Thief.
Mild Shock schrieb:
Hi,
Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and
he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel
simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:
Parallel -C-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a
Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?
Bye
Ross Finlayson schrieb:
No, troll, these are serial algorithms their optimized forms.
Normal sorts of forms, ....
Yeah, everybody already figured out "interpreters" and
"programs" and "spawning".
Go spawn yourself.
Mild Shock schrieb:
Hi,
You are still chewing on SIMD. LoL
Ross Finlayson schrieb:
Then the idea is that any of those can be found and matched inless than
one "run", i.e. a stall-less, branch-less, call-less list of
a few or less than a few dozens or less than a few hundredsYou cannot make the mental translation that if you have:
instructions, the results "findings" in data and corresponding
"matchings" of expressions, that runs in less than one microsecond. >>>>>
Ross Finlayson schrieb:
So, the context then is for register state and stack contents, that >>>>> -a> the indicators of the above as "positive presence" then is to make >>>>> -a> for that the adjustments to the offsets and extents and the shifts >>>>> -a> is according to those, otherwise no-ops. Then the idea is that a >>>>>As independent logical thread state, that automatically MIMD follows? >>>>>
Whats the problem to solve then?
Bye
Mild Shock schrieb:
Hi,
Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:
Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein
He is also not Zweistein, since he doesn't
understand concepts such as:
- NVIDIA Volta ff. architecture
Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.
Bye
Mild Shock schrieb:
Hi,
Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world
country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.
The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:
1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI
Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its
keyboard and some telephathy module ?
Bye
Mild Shock schrieb:
Hi,
Ride the snake
He's old and his skin is cold
The west is the best
The west is the best
Get here and we'll do the rest
The blue bus is calling us
The blue bus is calling us
Driver, where you taking us?
Apocalypse Now intro: The Doors, The End {1979}
https://www.youtube.com/watch?v=CIrvSJwwJUE
Bye
Hi,
Again I posted everything here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
The repo says, same time when I posted
the link first time:
This repository was archived by the
owner on Jul 9, 2026. It is now read-only.
Now a USENET user, who had already entitled
himself for a couple of irrational accusations
towards my side, is asking this question:
Chris M. Thomasson schrieb, Jul 24, 2026
Show an outline of what you
need you compute shader to do?
Bravo, thats a delay of a wooping 15 days.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Rossy Boy was a generative AI before
the term existed. All his postes are huge
piles of copy pasta slop.
Not a single original thought, or even
some understanding what he writes. Nowadays
he uses Kimi to produce his copy pasta
slop. One result from his paper mill,
Even a bibliograph cannot help here.
Ross Finlayson:
RF rCo transcript received and read. The session closed well. What we mapped out across these rounds is, to my mind, a credible foundation:
two matcher normal forms (AND for properties, XOR for code-points), a validated SSE2 smearing sequence via Claude's S1/S2 sketch, a closed calling convention with explicit ABI spill gates, and the SBC-less
design mantra as a gradient rather than a boolean. The open items rCo stack tagging, AST wire format, bit-granular Viswath boundaries rCo are properly scoped for next time rather than lost.
"bit-granular Viswath boundaries" LoL
It probably refers to Rossy Boys "Wish he
knew What" he is talking about, Viswath is his
alter ego projection:
The unbounded gibber polymath.
Bye
Hi,
Now you can compare this here from 2008
with modern AI Laptops for 500-1000 USD:
Google spotlights data center inner workings https://web.archive.org/web/20131019063218/http://news.cnet.com/8301-10784_3-9955184-7.html
There is a striking similarity, only what
once occupied a rack, has now the size
of your plam, all inside one silicon chip:
- Multiple CPU cores on the same chip
- Multiple GPU units on the same chip
- Network on the same chip communication
- Crossbar caches on the same chip
- Disk controllers on the same chip
- Multi channel RAM access on the same chip
Pretty cool!
Bye
P.S.: Example such devices with iGPU:
Intel(R) Core(TM) Ultra 7 258V
AMD Ryzen AI 7 350 w/ Radeon 860M
Apple A18 Pro, Darwin Kernel Version 25.5.0
Snapdragon(R) X - X126100 - Qualcomm(R) Oryon(TM) CPU
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,..
They are the same:
Performance of the Cray T3D
https://arxiv.org/abs/hep-lat/9509003v1
GPU Backend: Find 0xCAFFEE with ?-WAM
https://medium.com/2989/8890efd3503c
The details of the power and cooling in he old Art Deco building
In comp.lang.prolog Mild Shock <janburse@fastmail.fm> wrote:
Hi,..
They are the same:
Performance of the Cray T3D
https://arxiv.org/abs/hep-lat/9509003v1
GPU Backend: Find 0xCAFFEE with ?-WAM
https://medium.com/2989/8890efd3503c
Moore backend.
I set up a small company ~2000 to build a supercomputer center
in my city.
The aim was to build a TF system for a couple mil compared with
big configs some banks and colleges were buying at the time for $200+ mn.
We had a schedule to optimize cost taking into account the doubling of
power expected every couple ys. Each month we'd buy another 30-60 boxes retail from local stores and clone and rack them up.
The details of the power and cooling in he old Art Deco building
we got cheap rent in can be left for another storytime.
Anyway... we got our system up, started getting a list of clients
together, making some money. Around 2004/5 a single-board GPU could
also do a TF (of a kind :) and was a bit cheaper than the $2mn we spent
on getting out 1100 boxes together.
Hi,
The details of the power and cooling in he old Art Deco buildingYes sure, Horsy Boy!
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
How it started, NVIDIA being cool:
NCCL provides routines such as all-gather,
all-reduce, broadcast, reduce, reduce-scatter,
and point-to-point send and receive. These
routines are optimized to achieve high
bandwidth and low latency over PCIe,
NVIDIA NVLinkrao, and other high-speed
interconnects within a node and over
NVIDIA networking across nodes.
https://developer.nvidia.com/nccl
How its going, vLLM trying to be cool:
[RFC]: Native Weight Syncing APIs
However, there are no standardized methods for
performing online weight syncing. Open source projects
like SkyRL, VeRL, and TRL need to include their
own implementations of the weight syncing
infrastructure, leading to added complexity
for developers seeking to adopt vLLM as their
inference server for post-training workloads. https://github.com/vllm-project/vllm/issues/31848
How much Workers are enough? I guess it depends
on I/O parallelism, CPU Memory parallelism, CPU
Processing parallelism, and now also
GPU Memory parallelism and GPU Processing
parallelism, and last but least you might have
a couple DMAs sitting here and there,
or even invoking a sort of RDMA. Quite amazing!
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:
RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data
But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.
So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.
But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:
In-Browser LLM Inference Engine
https://webllm.mlc.ai/
My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.
But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese
hardware such some new Kirin AI in 2027.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Have Fun!
Bye
Mild Shock schrieb:
On 8/4/26 05:42, Mild Shock wrote:
Have Fun!
Bye
Mild Shock schrieb:
I am thinking that mind uploading would require
advanced microscopy to read all of the logic
within the connections between the coding neurons
in the brain, the axons, dendrites, and synapses,
as well as massively parallel computing to simulate
the operation of such an uploaded brain in reasonable
amounts of time.
Then of course there is the series 'Upload'.
The moral of the story of course is that the secret
to immortal life is that you need to get a subscription
to Amazon Prime.
Hi,
Nice try Rossy Boy --> **plonk**
Bye
P.S.: Woa!
My killfile is growing, and growing...
x3 schrieb:
On 8/4/26 05:42, Mild Shock wrote:
Have Fun!
Bye
Mild Shock schrieb:
I am thinking that mind uploading would require
advanced microscopy to read all of the logic
within the connections between the coding neurons
in the brain, the axons, dendrites, and synapses,
as well as massively parallel computing to simulate
the operation of such an uploaded brain in reasonable
amounts of time.
Then of course there is the series 'Upload'.
The moral of the story of course is that the secret
to immortal life is that you need to get a subscription
to Amazon Prime.
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly. https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Recently there was a paper somebody mentioning
a flit doing a ACK or NACK, to express
backpressure inside a Network on a Chip.
But what is a flit? It seems multiple
flits can be used to create the message
passing in one directiob before the
ACK or NACK in the other direction?
"The growing need for performance from
computing systems drove the industry into
the multi-core and many-core arena. In this
setup, the execution of a kernel (a program)
is split across multiple processors and the
computation happens in parallel
Flits represent logical units of information,
while phits represent the physical domain,
that is, phits represent the number of bits
that can be transferred in parallel in a
single cycle. Consider the Cray T3D. It has
an interconnection network which uses
flit level message flow control wherein each
flit is composed of eight 16-bit phits. That
means its flit size is 128bits and phit size
is 16bits. Also consider the IBM SP2 switch.
It also uses the flit level message flow
control, but its flit size is equal to its
phit size, which is set to 8 bits." https://en.wikipedia.org/wiki/Flit_(computer_networking)#Example
Well my idea how this is realized in silicon
is rather foggy, I mean even the Hack project
from Nand 2 Tetris, does not show some gate level
schemes for flits and phits.
Could be an interesting extension. But somehow
the image of flits and phits inspired my channel
objects here below. But I am afraid they are fire
and forget, no ACK and NACK:
-C-WAM Contest: 1 Million Packets with Prolog https://medium.com/2989/ec3e91551773
Its amazing that a max_size(1) buffer
can beat an unbounded buffer!
LoL
Bye
Mild Shock schrieb:
Hi,
How it started, NVIDIA being cool:
NCCL provides routines such as all-gather,
all-reduce, broadcast, reduce, reduce-scatter,
and point-to-point send and receive. These
routines are optimized to achieve high
bandwidth and low latency over PCIe,
NVIDIA NVLinkrao, and other high-speed
interconnects within a node and over
NVIDIA networking across nodes.
https://developer.nvidia.com/nccl
How its going, vLLM trying to be cool:
[RFC]: Native Weight Syncing APIs
However, there are no standardized methods for
performing online weight syncing. Open source projects
like SkyRL, VeRL, and TRL need to include their
own implementations of the weight syncing
infrastructure, leading to added complexity
for developers seeking to adopt vLLM as their
inference server for post-training workloads.
https://github.com/vllm-project/vllm/issues/31848
How much Workers are enough? I guess it depends
on I/O parallelism, CPU Memory parallelism, CPU
Processing parallelism, and now also
GPU Memory parallelism and GPU Processing
parallelism, and last but least you might have
a couple DMAs sitting here and there,
or even invoking a sort of RDMA. Quite amazing!
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Now I implemented some multiple producer
and multiple consumer channel objects for
WebGPU. The only API to integrate it user
facing into pi-WAM is this single predicate:
/**
-a* flit(C):
-a* The predicate succeeds in C with a new channel. The channel
-a* can be used from within GPU backed -C-WAM logical threads.
-a*/
The Mac Neo is a Budget Monster. While the
Ryzen AI Laptop cost around 1300.- CHF.
The Mac Neo was around 600.- CHF with all
extras. Here some performance results,
checking out whether channel objects scale,
when increasing their number to
communicate the same 1 millon packets:
Java performance:
AI Laptop-a-a-a Single-a-a-a Double
Ryzen-a-a-a 705.1-a-a-a 337.4
Neo-a-a-a 669.4-a-a-a 239.9
WebGPU performance:
AI Laptop-a-a-a Single-a-a-a Double
Ryzen-a-a-a 731.8-a-a-a 392.9
Neo-a-a-a 932.8-a-a-a 483.5
Cool! Java is also pretty cool, their
semaphore library is top notch. I couldn't
replicate the resulst with JavaScript yet,
seems their Atomics.wait() resp. Atomics.waitAsync()
is totally broken, using futex is mutex for
fools somehow. I also found some gremlins
attacking one of the GPUs. The Intel AI Laptop
fails the above experiment. Maybe its a driver
Vulkan versus OpenCL or something problem,
or the Lunar lake architecture is nonsense.
Bye
Mild Shock schrieb:
Hi,
Recently there was a paper somebody mentioning
a flit doing a ACK or NACK, to express
backpressure inside a Network on a Chip.
But what is a flit? It seems multiple
flits can be used to create the message
passing in one directiob before the
ACK or NACK in the other direction?
"The growing need for performance from
computing systems drove the industry into
the multi-core and many-core arena. In this
setup, the execution of a kernel (a program)
is split across multiple processors and the
computation happens in parallel
Flits represent logical units of information,
while phits represent the physical domain,
that is, phits represent the number of bits
that can be transferred in parallel in a
single cycle. Consider the Cray T3D. It has
an interconnection network which uses
flit level message flow control wherein each
flit is composed of eight 16-bit phits. That
means its flit size is 128bits and phit size
is 16bits. Also consider the IBM SP2 switch.
It also uses the flit level message flow
control, but its flit size is equal to its
phit size, which is set to 8 bits."
https://en.wikipedia.org/wiki/Flit_(computer_networking)#Example
Well my idea how this is realized in silicon
is rather foggy, I mean even the Hack project
from Nand 2 Tetris, does not show some gate level
schemes for flits and phits.
Could be an interesting extension. But somehow
the image of flits and phits inspired my channel
objects here below. But I am afraid they are fire
and forget, no ACK and NACK:
-C-WAM Contest: 1 Million Packets with Prolog
https://medium.com/2989/ec3e91551773
Its amazing that a max_size(1) buffer
can beat an unbounded buffer!
LoL
Bye
Mild Shock schrieb:
Hi,
How it started, NVIDIA being cool:
NCCL provides routines such as all-gather,
all-reduce, broadcast, reduce, reduce-scatter,
and point-to-point send and receive. These
routines are optimized to achieve high
bandwidth and low latency over PCIe,
NVIDIA NVLinkrao, and other high-speed
interconnects within a node and over
NVIDIA networking across nodes.
https://developer.nvidia.com/nccl
How its going, vLLM trying to be cool:
[RFC]: Native Weight Syncing APIs
However, there are no standardized methods for
performing online weight syncing. Open source projects
like SkyRL, VeRL, and TRL need to include their
own implementations of the weight syncing
infrastructure, leading to added complexity
for developers seeking to adopt vLLM as their
inference server for post-training workloads.
https://github.com/vllm-project/vllm/issues/31848
How much Workers are enough? I guess it depends
on I/O parallelism, CPU Memory parallelism, CPU
Processing parallelism, and now also
GPU Memory parallelism and GPU Processing
parallelism, and last but least you might have
a couple DMAs sitting here and there,
or even invoking a sort of RDMA. Quite amazing!
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
:- multifile(strings/3).
/* de = ISO locale atoms with prefix de_ */ strings('evaluation_error.zero_divisor', de, 'Nulldivision.').
/* '' = fall back ISO locale atoms */ strings('evaluation_error.zero_divisor', '', 'Division by zero.').
Hi,
But, I still don't know what you main goal is?
It explicity says "Prolog inferencing" in
this phrase:
shave off some of the TOPS to do Prolog inferencing
It nowhere says draw some fancy stuff into
a Web canvas.
Bye
Mild Shock schrieb:
Hi,
It has textures to work with in the pipeline.
Hi,
Why would I use text inside my compute shader.
Could you tell me. The Hack VM doesn't do
textures. You are confused. There is nothing
about textures here:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
You can read the text , it says nowhere
consume or produce textures. Its not a rendering
application. I use the compute shader to run Prolog:
"At the end of 2025 we acquired a couple of
AI Laptops , that were still cheap, since
RAM prices had not yet rocketed. The intend
was to tap into the Copilot+ certified hardware,
and shave off some of the TOPS to do Prolog
inferencing. Amazingly our -C-WAM can
churn 11.4 GIGA LIPS.
GPUs have evolved form lock-step to independent
thread scheduling. This made it possible to
port the Hack VM variant, that forms the basis
for our -C-WAM, to WebGPU computer shaders.
Using NUM_SHADERS = 4096 we could produce
11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M."
Bye
Mild Shock schrieb:
Hi,
WebGPU and WebGL are two different things. I explained
that towards you already like 3-5 times.
Of course we can make a special texture to handle it.
You still don't understand that I am using WebGPU,
and not WebGL. WebGPU has three improvements,
that from your talking are missing in WebGL?
- It has compute shaders
- It has arrays
- It has structs
- What else?
I didn't use structs in my example, although Gemini
nearly forced me to use structs. But you could
use a struct with fields and some of these arrays
to represent a queue. But here in this example
that is open source, I only used flat arrays. I
nowhere needed to abuse textures to store something:
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
You can study the source code, the arrays have
CUDA inspired binding annotations but are not
CUDA but rather WGSL:
Hack VM as a Compute Shader in WGSL
@group(0) @binding(0) var<storage, read> code: array<i32>;
@group(0) @binding(1) var<storage, read_write> state: array<i32>;
https://github.com/Jean-Luc-Picard-2021/gigabudget/blob/main/course/example63/boot.mjs
You can say whether a buffer is read, or read_write.
Buffers can be transfered from CPU to GPU, before
running commands, and transfered back from GPU to
CPU after running commands. The use case that
you find on GitHub uses both. Namely also fetching
results via a buffer, to then show them in
the HTML page. Shouldn't be much a problem to
run the example at home locally, all you need is
a HTTPS server. But the example is not yet queues,
but it already shows the foundation, which is WebGPU
with its language WGSL and not WebGL with its language
GLSL. These are two different things.
I explained that towards you already like 3-5 times.
Bye
Chris M. Thomasson schrieb:
On 8/1/2026 5:22 AM, Mild Shock wrote:shader. Of course we can make a special texture to handle it. But, we
Hi,[...]
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You are going to need a place to allocate nodes in the compute
need to strive to avoid a wait condition. I don't want a compute
shader to spin. Yes, CAS can be used, but, try to make it be used as
a "state machine", where the transitions from states are atomic. Try
to avoid it making a loop, where we loop on failure.
Mild Shock schrieb:
Hi,
As easy as queues and FIFO objects might
sound. They don't like congestion. NACK for
retransmission might double the Manhattan Distance:
You have not only start
S to end E communication:
+----E
|
|
S
You might also have ACK or NACK
from E or midpoints back to S:
-a-a-a S'
-a-a +
-a-a+
E'
Ok, I made that up, I have no idea what a flit is,
when the author wrote this here:
"Packet flits are held in the FIFO which can
be used to determine back pressure. Dropping flits
in a NoC may not be possible since these
architectures may not provide an end-to-end
protocol for retransmission."
Routing Algorithms for 2D NoC Architectures
http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
Mild Shock schrieb:
Hi,
Looking at the floor plan of a NPU:
Getting peak TOPS on a Ryzen AI 7 350 NPU
https://destevez.net/2026/05/getting-peak-tops-on-a-ryzen-ai-7-350-npu/ >>>>>
It seems to me comms between tiles takes
at least Manhattan Distance or L1 Norm time,
if there is no comms congestion
But how does a packet travel? This way:
+----E
|
|
S
Or this way, from start S to end E:
-a-a-a +-E
-a-a +
-a-a+
S
And what does the chip do if there is
traffic congestion? Some papers are
here, possibly an old problem giving
that processor "cubes" are nothing new.
But a "cube" would be 3D and not 2D.
This paper is old from 2007 or so:
Routing Algorithms for 2D NoC Architectures
http://cva.stanford.edu/classes/ee382c/research/2DRouting.pdf
Bye
Hi,
Why is nobody mentioning Agda here. It has
beautiful dependent types, and tactics are
just programs. Poor Henk Barendregt, not
everybody likes dependent types it seems:
Are we stuck with Lean?
https://mathoverflow.net/q/513742/
Does Depependent types require proof objects,
which waste large amounts of memory. Well,
if you are not good in erasing them.
But is there a Red Pyjama for Proof Assistants,
the baby cradle where LLMs can learn proof
assistant lingua and strategies. It seems
yes, synthetic data corpuses to the rescue:
We address this gap by introducing SMAD
(Synthetic Multilanguage Autoformalization
Dataset), a 400K 4-to-3 parallel corpus
covering four formal languages (Dedukti,
Agda, Coq, Lean) and three natural languages (
English, French, Swedish), generated via
the Informath project.
https://github.com/GrammaticalFramework/informath
But the corpus could be an accident, maybe rather
a toy from the https://www.grammaticalframework.org/
folks, will this have an impact?
Bye
Mild Shock schrieb:
Hi,
Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:
RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data
But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.
So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.
But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:
In-Browser LLM Inference Engine
https://webllm.mlc.ai/
My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.
But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese
hardware such some new Kirin AI in 2027.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
Hi,
Years ago Sam Altman said to have no idea how
to generate revenue, but when the generally
intelligent system is in place, he might ask it.
Some schools approach the rCLgeneralityrCY from
a totally wrong perspective. Take the EyeProlog
Pseudo Scientism here:
The Art of EyeProlog https://eyereasoner.github.io/eyeprolog/the-art-of-eyeprolog
It is the same nonsense like constraint propagation,
the idea here is to evolve better software, that it
has as a main component refinement:
Start -> Algo1 -> Algo2 -> Algo3 -> Algo4 ...
But EyeProlog itself is an example of not using
this refinement. Like dropping the classical
WAM architecture, and back to YieldProlog somehow.
What if the world ticks like this
when it come to generality:
-a-a-a-a-a-a /-> Algo1
-a-a-a-a-a /--> Algo2
Start ---> Algo3
-a-a-a-a-a \--> Algo4
-a-a-a-a-a-a \-> ...
Innovation requires to start from scratch.
I think this little booklet, recommended by
Ernst Specker, Proofs from THE BOOK is a
book of mathematical proofs by Martin Aigner
and G|+nter M. Ziegler, first published in 1998.
Just wants to teach us about this bifurcation:
Chapter 1: Six proofs of the infinity of
the primes, including Euclid's and Furstenberg's. https://en.wikipedia.org/wiki/Proofs_from_THE_BOOK
Yeah, lets aim for surprises by
generative AI, not refinement.
Bye
See also:
Sam Altman on his Business Model
https://www.youtube.com/shorts/pLnyjxgFxew
Mild Shock schrieb:
Hi,
Why is nobody mentioning Agda here. It has
beautiful dependent types, and tactics are
just programs. Poor Henk Barendregt, not
everybody likes dependent types it seems:
Are we stuck with Lean?
https://mathoverflow.net/q/513742/
Does Depependent types require proof objects,
which waste large amounts of memory. Well,
if you are not good in erasing them.
But is there a Red Pyjama for Proof Assistants,
the baby cradle where LLMs can learn proof
assistant lingua and strategies. It seems
yes, synthetic data corpuses to the rescue:
We address this gap by introducing SMAD
(Synthetic Multilanguage Autoformalization
Dataset), a 400K 4-to-3 parallel corpus
covering four formal languages (Dedukti,
Agda, Coq, Lean) and three natural languages (
English, French, Swedish), generated via
the Informath project.
https://github.com/GrammaticalFramework/informath
But the corpus could be an accident, maybe rather
a toy from the https://www.grammaticalframework.org/
folks, will this have an impact?
Bye
Mild Shock schrieb:
Hi,
Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:
RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data
But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.
So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.
But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:
In-Browser LLM Inference Engine
https://webllm.mlc.ai/
My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.
But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese
hardware such some new Kirin AI in 2027.
Bye
Mild Shock schrieb:
Hi,
Remember when first all local AI was Python
and PyTorch APIs. And then suddently people started
using bare metal C/C++ Code. Here is the story:
How it started:
GPT-J or GPT-J-6B is an open-source large
language model (LLM) developed by EleutherAI
in 2021. As the name suggests, it is a
generative pre-trained transformer model
designed to produce human-like text that
continues from a prompt.
https://www.eleuther.ai/
How it was going [Georgi Gerganov]:
So a few days later comes out the LLaMA, I do
some calculations and I figure out rCLOkay, 65
billion parameters. You probably need about
40 gigs of RAM, with 4-bit quantization. So
this can run on a MacBook. Why not do it?rCY
Why I was able to do it so quickly - basically,
for all that I saw itrCOs pretty much GPT-J architecture
with some modifications, like some extra memorization
layers. ItrCOs minor changes. Basically, again, the
existing code for the GPT-J, I just simply
modified it there, it happened pretty quickly.
https://changelog.com/podcast/532
Georgi Gerganov, Bulgarian, now with Hugging
Face, ggml-cann also running on Chinese AI chips.
ggml Manifesto https://github.com/ggml-org/ggml
Bye
| Sysop: | Amessyroom |
|---|---|
| Location: | Fayetteville, NC |
| Users: | 74 |
| Nodes: | 6 (0 / 6) |
| Uptime: | 46:34:19 |
| Calls: | 1,100 |
| Files: | 1,339 |
| Messages: | 275,493 |