Lianmin Zheng
|
463baafe10
|
[Auto Sync] Update batch_invariant_ops.py (20260221) (#19098)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Jiayi Yuan <34369239+jy-yuan@users.noreply.github.com>
|
2026-02-20 18:27:31 -08:00 |
|
Yisheng Gong
|
1c4616a034
|
fix: add bias when enable mm fallback variant (#17690)
|
2026-01-28 09:50:49 +08:00 |
|
fzyzcjy
|
2fbc78a083
|
Support fast gemm when in batch invariant DeepGEMM fallback (#13259)
|
2025-11-15 16:34:15 +08:00 |
|
Minglei Zhu
|
8a4373405e
|
re-submit 12911 but relax the requirement for deepgemm (#13226)
|
2025-11-15 15:37:12 +08:00 |
|
fzyzcjy
|
86255f27b4
|
Revert "fallback to triton mm_persistent kernel when deepGemm fail" (#13178)
|
2025-11-12 22:03:35 -08:00 |
|
Lianmin Zheng
|
9ea2c686c7
|
[Auto Sync] Update batch_invariant_ops.py (20251109) (#12916)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
|
2025-11-10 01:51:39 -08:00 |
|
Minglei Zhu
|
8a821af793
|
fallback to triton mm_persistent kernel when deepGemm fail (#12911)
|
2025-11-08 23:42:26 -08:00 |
|
Minglei Zhu
|
c14cc47e39
|
[Deterministic] Optimize bmm_batch_invariant op (#12522)
|
2025-11-04 00:33:31 -08:00 |
|
Yuzhen Zhou
|
0380ca82ef
|
Add Batch‑Invariant RMSNorm (#12144)
|
2025-10-28 21:05:57 -07:00 |
|
fzyzcjy
|
0103f374ba
|
Support DeepGEMM for deterministic inference (#12142)
|
2025-10-26 22:36:17 +08:00 |
|
fzyzcjy
|
c001deba37
|
Make bmm batch invariant injection optional (#12118)
|
2025-10-26 10:18:35 +08:00 |
|
Minglei Zhu
|
f4b78d137c
|
[1/2] deepseek deterministic: support deterministic inference for deepseek arch models on a single GPU (#12000)
|
2025-10-24 15:17:28 -07:00 |
|
Stefan He
|
eae9a9fb9d
|
Fix batch invariant ops (#11368)
|
2025-10-10 20:49:08 -07:00 |
|
Stefan He
|
86527a4799
|
[deterministic inference] Move batch invariant pkg to sglang (#10695)
|
2025-09-21 19:35:14 -07:00 |
|