Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

X86 optimized doesn't actually mean that much on a lot of workloads because the memory hierarchy completely dominates performance.

Apple knew this so the M1 has an extremely fast memory system.



It’s a scientific code which manipulates a lot of relatively big matrices, which saturates the FPUs pretty well (I made the analysis back then).

Intel’s memory architecture can keep up with the load until I utilize all cores. When I start to use HT cores, neither cache, nor memory can keep up.

I do the both tests (Intel and Apple) both with physical core counts to keep it fair.

So it’s not some run of the mill code.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: