Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Can we start using stainless steel instead?by amelius
- Hot take: there is no portable SIMD.
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
by Archit3ch - I am using my own lib, `lin_alg`, which apes core_simd for floating point values, and extends the concept to vectors and quaternions. I will eventually replace the floating point portions with core::simd upon its arrival in stable Rust.
Downside: It's currently x86 only.
- "Hardware that does arithmetic is cheap, so any CPU made this century has plenty of it. But you still only have one instruction decoding block and it is hard to get it to go fast, so the arithmetic hardware is vastly underutilized.
Warning, Nitpick. Saying "hardware [...] is cheap" and "instruction decoding is expensive" (implied) is a minor contradiction. You can actually duplicate instruction decoding just fine, it's the thing before it where things go to hell: Instruction fetching.
Nothing prevents you from building a computer that can fetch, decode and execute 16 instructions at once, assuming they don't all write to the same register.
But if you want to do the same add repeated 16 times you'll need 16 times more program memory and 16 wider read ports on your caches and so on. SRAM is really expensive so this strategy will waste a lot of area on memory that you probably didn't need in the first place. I say this as someone who had to design a chip in university and basically you couldn't even find the primitive CPU in-between the massive SRAM blocks. By reusing the same instruction you can now increase your compute to memory ratio in terms of area.
Just a heads up for people who want to know why SIMD is a thing. I'm not criticizing the article, I just want more people to realize the pain that SRAM represents to chip designers.
by imtringued - I am hoping for portable SIMD so much. But I still think that often a manually rolled SIMD will be faster.
Also the state of SIMD in Cranelift is also very WIP. They pretty much just support a subset of 128bit vectors with some rare exceptions.
by sharktheone - > ... you can just assert that they’re all recent enough to at least have AVX2 that was introduced over 10 years ago, and have the program crash or misbehave if it ever runs on anything without AVX2
> However, if you are distributing the binaries for other people to run, that’s not really an option.
This all depends on what kind of software you're making. A lot of games set their requirements about 5 generations back, like FC 27 where the minimum is a Ryzen 1600. That lets them use AVX2 unconditionally and prevent complaints from users who tried to run it with a super old CPU.
Then you get whole Linux distros like CachyOS and Clear (RIP) that rebuild the world for each architecture level and have them as separate variants. I think it still counts as binaries for other people.
by tancop - AArch64 definitely has a much more comprehensive baseline than x86-64, but there are some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions. And unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting which intrinsics require specific FEAT_* flags.
The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.
by ack_complete