Rendered at 19:19:20 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
AlanZucconi 3 hours ago [-]
I'm really curious... How did you manage to post this link before I did?
tobr 3 hours ago [-]
I have your feed in my RSS reader! This seemed like something HN would be interested in.
AlanZucconi 2 hours ago [-]
I should be faster next time then: my entry got tagged as [dupe] ahah!
Also: I somehow got 10x the usual amount of traffic today, and my website is sort of on fire!
> No multiplications. No divisions. No lookup tables. Just a few bitwise instructions.
Would have appreciated this article more if it was written by a human.
AlanZucconi 1 minutes ago [-]
I understand that a big portion of online content is now 100% AI-generated, and that can be somewhat problematic. But for creators like me, who have been publishing articles and books for over 10 years, this AI witch hunt can be quite demotivating.
I worked on this project for over one year. I wrote an entire distributed framework to calculate maximal triplets, and I have 130+ machines running 24/7 for 12 weeks on N=8192. This article is an extended version of the script for the video documentary that will be released before the end of the year.
If you look back at my website, I used to publish two small articles a week. I've since reduced to 1 or 2 large pieces a year. And one of the reasons was exactly to rise above the many blogs that post small, fragmented articles, which could be generated in 2 minutes by ChatGPT. If I wanted to continue in that direction, I could be publishing 100 short articles a week with ChatGPT.
I started writing practical shader tutorials back in 2014 because there were not enough good, accessible resources online. I'm now focusing on large, in-depth pieces with original research, because that's what's valuable right now that ChatGPT has replaced StackOverflow.
If you appreciated this article, I hope you'll appreciate it even more knowing that YES, it was written by a human (it's me, hi!). <3
I’m not sure what point you’re trying to make, exactly, but a use case for better non-CS generators has always been stochastic simulation, especially simulation/sampling approaches that are bound by the number and quality of uniform variates per second.
As someone who has spent considerable time working in these areas, I still appreciate advances.
saithound 2 hours ago [-]
> I’m not sure what point you’re trying to make,
Have you skimmed the linked thread?
> especially simulation/sampling approaches that are bound by the number and quality of uniform variates per second
Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.
But even if you insist that you alone are doing some very weird stochastic simulation which is somehow bottlenecked on sourcing random numbers fast enough, the falling in planes phenomenon linked above would make xorshift-type generators a poor choice for most sorts of simulations. It introduces spatial correlations into any sort of lattice dynamics simulation (Ising model, percolation) and every high dimensional Monte Carlo integration. Beyond falling in the planes, since xorshift is linear over GF(2), it is also a particularly bad choice for nondeterministic cellular automata and Boolean dynamical systems which use parity, bit masks, or xors.
AES-CTR throughput on a modern CPU is higher than that of xoshiro256++, and much higher quality. No advances in non-CS PRNGs can beat that while maintaining the same quality. If your stochastic simulation is bottlenecked on random bits, CSPRNGs are still the way to go, and they don't interact in nasty ways with any dynamical system you can actually sinulate quickly.
moregrist 54 minutes ago [-]
> Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.
Actually, I spent a considerable amount of time in my doctorate and postdoc doing this.
Any kind of MCMC sampling of a simple model tends to be bound by the rate you can draw variates.
Examples of this include: Gillespie simulations of chemical kinetics, Ising and Potts lattice models (including their roughly bazillion variations), and anything resembling bootstrap or permutation sampling.
Just because your problems aren’t bound by the rate of drawing uniform variates doesn’t mean that these problems don’t exist. It just means that you have a narrow view.
dgacmu 2 hours ago [-]
This is a very weird hill to die on.
I do a lot of testing and designing of things like hash tables and filters, and having a really fast, non-CS generator is incredibly useful for being able to clearly identify performance bottlenecks in designs. PCG has been spectacularly useful for that purpose for me.
saithound 2 hours ago [-]
When was the last time a new PRNG helped you clearly identify a performance bottleneck?
As in, you were using state of the art generator X, and you couldn't see the performance bottleneck, but updating to a newer (faster, or same speed but higher quality) generator Y, and could subsequently identify the performance bottleneck?
If you're using PCG, not in the last 12 years.
(In a parallel comment I suggest trying AES-CTR for this use case)
dgacmu 2 hours ago [-]
It's not critical but if you gave me something that behaved statistically like PCG (i.e., I didn't fret about whether it was going to cause me weird problems) but was twice as fast I'd be happy and would shift to it - it would speed up profiling and measuring and that would be nice. We still find ourselves often pre-generating a list into memory to keep the prng entirely off of the measurement path. It wouldn't be magic, but I don't need magic. I like nice things that make my life a little easier in a small corner of my research. :)
Straw 51 minutes ago [-]
Most modern RNGs should be faster than memory bandwidth (when optimized), so unless your list is small enough to fit in cache, its unclear if this is faster?
saithound 2 hours ago [-]
dgacmu: if you're writing C on x64, try AES-128-CTR (AES-NI, 8 way) using the header wmmintrin.h which has hardware accelerated primitives for this. An LLM can implement the RNG for you based on this comment if you want to test it out quickly. It should be faster than PCG, and higher quality.
Would have appreciated this article more if it was written by a human.
I worked on this project for over one year. I wrote an entire distributed framework to calculate maximal triplets, and I have 130+ machines running 24/7 for 12 weeks on N=8192. This article is an extended version of the script for the video documentary that will be released before the end of the year.
If you look back at my website, I used to publish two small articles a week. I've since reduced to 1 or 2 large pieces a year. And one of the reasons was exactly to rise above the many blogs that post small, fragmented articles, which could be generated in 2 minutes by ChatGPT. If I wanted to continue in that direction, I could be publishing 100 short articles a week with ChatGPT.
I started writing practical shader tutorials back in 2014 because there were not enough good, accessible resources online. I'm now focusing on large, in-depth pieces with original research, because that's what's valuable right now that ChatGPT has replaced StackOverflow.
If you appreciated this article, I hope you'll appreciate it even more knowing that YES, it was written by a human (it's me, hi!). <3
There is no real use case for better non-CS generators, as explained by adrian_b back in 2021 [3].
[1] https://en.wikipedia.org/wiki/RANDU [2] https://arxiv.org/abs/1908.10020 [3] https://news.ycombinator.com/item?id=28886698
As someone who has spent considerable time working in these areas, I still appreciate advances.
Have you skimmed the linked thread?
> especially simulation/sampling approaches that are bound by the number and quality of uniform variates per second
Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.
But even if you insist that you alone are doing some very weird stochastic simulation which is somehow bottlenecked on sourcing random numbers fast enough, the falling in planes phenomenon linked above would make xorshift-type generators a poor choice for most sorts of simulations. It introduces spatial correlations into any sort of lattice dynamics simulation (Ising model, percolation) and every high dimensional Monte Carlo integration. Beyond falling in the planes, since xorshift is linear over GF(2), it is also a particularly bad choice for nondeterministic cellular automata and Boolean dynamical systems which use parity, bit masks, or xors.
AES-CTR throughput on a modern CPU is higher than that of xoshiro256++, and much higher quality. No advances in non-CS PRNGs can beat that while maintaining the same quality. If your stochastic simulation is bottlenecked on random bits, CSPRNGs are still the way to go, and they don't interact in nasty ways with any dynamical system you can actually sinulate quickly.
Actually, I spent a considerable amount of time in my doctorate and postdoc doing this.
Any kind of MCMC sampling of a simple model tends to be bound by the rate you can draw variates.
Examples of this include: Gillespie simulations of chemical kinetics, Ising and Potts lattice models (including their roughly bazillion variations), and anything resembling bootstrap or permutation sampling.
Just because your problems aren’t bound by the rate of drawing uniform variates doesn’t mean that these problems don’t exist. It just means that you have a narrow view.
I do a lot of testing and designing of things like hash tables and filters, and having a really fast, non-CS generator is incredibly useful for being able to clearly identify performance bottlenecks in designs. PCG has been spectacularly useful for that purpose for me.
As in, you were using state of the art generator X, and you couldn't see the performance bottleneck, but updating to a newer (faster, or same speed but higher quality) generator Y, and could subsequently identify the performance bottleneck?
If you're using PCG, not in the last 12 years.
(In a parallel comment I suggest trying AES-CTR for this use case)