Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Yes, but the contended case is often the critical part of an application’s performance. Weakly ordered CPUs therefore need fast barrier mechanisms, which in turn means re-creating much of the store-ordering, writeback buffering, and commit logic anyway.
  • I'd be more concerned with instruction reordering by the compiler. So even if the hardware guarantees some ordering, you will still need those memory fences.
  • The good part is that at least there's some explicit C++ std::memory_order and std::sync::atomic::Ordering lets me pick where it really matters.

    For example, I built a skew handling model which needed low overhead cross-thread counters, where Ampere and Graviton was different from the M1 mac in benchmark - even down to the same assembly on different systems (cmov specifically).

  • Would be useful for the article to start with a word on what memory ordering is.

    Edit: it means, (re)ordering of memory access operations esp. when multiple execution units (cores) operate in parallel. Weakly ordered (ARM, RISC-V): if Core 1 writes Memory 1 then Memory 2, but for some reason writing into Memory 2 is quicker, then Core 2 may see Memory2 written before Memory 1 is written. While strong ordering (x86) retains the original order.

Explore Birbla archives

Memory Ordering in CPUs · Birbla