A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation

Global Tech Moderate confidence — 64/100
Sources: Arxiv