Gabriel Wu
|
bfe983c4c2
|
Refactor JIT compilation (+NVRTC support) (#94)
* [wip] refactor: compile to .cubin
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* refactor: compile to .cubin and add NVRTC option
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* fix: compiler version
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* feat: compat for old drivers
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* feat: save kernel name to file
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* feat: fix win compat
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* fix: windows compat
Signed-off-by: Gabriel Wu <13583761+lucifer1004@users.noreply.github.com>
* feat: make API more general
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* feat: drop support for CUDA<12.3
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* doc: update README
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
* Some lints and refactor
* Refactor runtime
* Several fixes
* Refactor environment variables
* Code format
* Add a TODO
* Compatible with CUDA 12.3
* Fix indent
* Fix typing
* Drop support for Windows
* Add a TODO
---------
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: Gabriel Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Chenggang Zhao <chenggangz@deepseek.com>
|
2025-05-07 11:38:14 +08:00 |
|
yukuai26
|
95e81b3dd6
|
Indivisible TMA (#90)
Fix indivisible shapes for TMA multicast
---------
Co-authored-by: yukuai <yukuai@deepseek.com>
Co-authored-by: Chenggang Zhao <chenggangz@deepseek.com>
|
2025-04-23 14:55:14 +08:00 |
|
yukuai26
|
891f35adf5
|
Support TMA multicast on B with m_grouped_gemm_contiguous. (#88)
|
2025-04-21 09:43:17 +08:00 |
|
Chenggang Zhao
|
7ffb118e54
|
Support multicasting on B
|
2025-03-25 14:56:42 +08:00 |
|
Chenggang Zhao
|
b922e64cb2
|
Support block size 160
|
2025-03-25 13:37:59 +08:00 |
|
sazc
|
46eb0d08fb
|
Performance: Larger BlockTile optimizations enable 1470+ TFLOPS FP8 performance on the H800-SXM platform
|
2025-03-25 10:44:57 +08:00 |
|
AcraeaTerpsicore
|
96b31fd6bb
|
fix typo
|
2025-02-26 18:37:22 +08:00 |
|
Chenggang Zhao
|
a6d97a1c1b
|
Initial commit
|
2025-02-25 22:52:41 +08:00 |
|