Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

    Explore Birbla archives

    A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation · Birbla