I'm an old dog trying to learn a new trick. Old dog says Nvidia GPU with 8+GB of RAM and thousands of cores is powerful. New trick says M1 with max 8 cores and ??? amount of RAM is more powerful??? My head hurts trying to come up with how. Where is the magic happening that makes 6K+ video run in real time in Resolve, when an Intel CPU with multiple GPUs needs a good wind at its back going downhill. Of the many hand wavy demo details left out, what kind of video are they using? MP4 video, RAW videos, ProRes, etc? Is Resolve running in realtime just playing back the video but immediately chokes when you apply a single node with light grade applied? The M1 release videos and too PR speak for me
• NVIDIA is on a dumpster fire cheap “Samsung 8nm” process that is quite possibly the worst <=14nm process especially when it comes to power and heat.
• Apple has the benefits of complete vertical integration, both on the hardware and software side.
• Neural engine is essentially Tensor Cores in NVIDIA’s GPU but occupies at least 4x the equivalent die area as tensor cores (no public details on performance
yet).
• NVIDIA doesn’t want to make their consumer GPUs too powerful on tensor operations in order to not cannibalise their 1000% markup ML cards.
It’s almost a classic Intel: financial greed and financial engineering, combined with complacency from being long for so long.
Heck, AMD’s new top end card is tied with the 3090 - but $500 cheaper.
NVIDIA is reportedly scrambling to try and get back into TSMC who is going to make an example out of them.
Wait, you're comparing M1's neural engine with Ampere's tensorcores? Apple is talking TOPS, not TFLOPS, that means M1's 11 TOPs are for inference only, not training. By contrast, a tensorcore on Ampere does ~20 TFLOPS FP64, 78 TFLOPS FP16/BF16, and with GPU boost clock, up to 312 TFLOPs @ BF16.
If you want to talk TOPS, an A100 does up to 1.2 exa-ops INT8, and 2.4 exa-tops INT4. That is, an A100 is more than 1000 times more powerful than the M1 at inferencing, while also supporting up to FP32/FP64 weights, and a 3080 is more than 500 times more powerful than the M1 neural engine.
Even an RTC 2060 beats it, the 2060 has 52 TFLOPs worth of tensor core ops (not counting the CUDA cores), and it has >100 TOPs @ INT8, or 200+ TOPs @ INT4, so more than 10x the M1.
It's actually very forgiving. For high frequency signal traces on a PCB there is this rule of thumb that your trace lengths must not differ by more 15cm per ns because that is how far electrons travel in one nanosecond. If you have a 10GHz signal that means you have a 0.01ns budget which translates to a 1.5mm max length difference between traces. Matching within 1.5mm is easy to do by hand.
Now lets talk about actual processors instead of doing an analogy. 15cm per ns is more than enough to travel everywhere inside your CPU but the vast majority of logic is localized (usually signals stay in the same core). It's only awful if you go off package to DRAM or a second socket but then the budget is often higher than 1ns. Apple probably scores a lot of performance points here because the RAM is so close to the CPU.
It's not. The limit is how fast you can switch transistors. Electrons can go a pretty long distance inside a clock cycle, and waiting multiple clock cycles for signals to go from core to core is normal and easy.
Technically electrons barely move at all (on clock-cycle timescales); it's the electric field that moves at ~2/3 speed of light. Compare eg, wind speed (particles) versus the speed of sound (field).