Indivisible TMA (#90)
Fix indivisible shapes for TMA multicast --------- Co-authored-by: yukuai <yukuai@deepseek.com> Co-authored-by: Chenggang Zhao <chenggangz@deepseek.com>
This commit is contained in:
co-authored by
yukuai
Chenggang Zhao
parent
891f35adf5
commit
95e81b3dd6
@@ -16,7 +16,7 @@ Despite its lightweight design, DeepGEMM's performance matches or exceeds expert
|
||||
- [x] Shared memory swizzling for output
|
||||
- [ ] Larger block size on N (up to 256)
|
||||
- [x] MoE scheduler with TMA multicast compatibility
|
||||
- [ ] Fix TMA multicast compatibility for indivisible shapes
|
||||
- [x] Fix TMA multicast compatibility for indivisible shapes
|
||||
- [ ] Weight gradient kernels for dense models
|
||||
- [ ] Weight gradient kernels for MoE models
|
||||
- [ ] Utility kernels for MoE models (as a pre-built CUDA library)
|
||||
|
||||
Reference in New Issue
Block a user