Xiaoyu Zhang
|
50691d7b49
|
[opt kimi k2 2/n] apply kimi k2 thinking moe_fused_gate (#13332)
|
2025-11-16 21:20:08 +08:00 |
|
 
|
aead0ef5e5
|
[FEAT][ROCM] enable fused shared expert for Rocm (#12201)
Co-authored-by: ZLkanyo009 <4071250045@qq.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2025-11-13 00:41:40 -08:00 |
|
Xiaoyu Zhang
|
a1cb717d0b
|
Opt kimi_k2_thinking biased topk module (#13150)
|
2025-11-12 16:52:06 -08:00 |
|
Ke Bao
|
58b12ccb46
|
Support piecewise cuda graph for deepseek v3 (#12996)
|
2025-11-10 23:18:03 +08:00 |
|
Nicolas Castet
|
2340798353
|
Register allgather/reducescatter buffers with symm memory (#12572)
|
2025-11-04 17:11:36 -08:00 |
|
Trevor Morris
|
dbcf85b7f0
|
Add --speculative-moe-runner-backend server arg (#10183)
|
2025-11-04 00:20:56 -08:00 |
|
Even Zhou
|
cafebef154
|
[NPU] bugfix for Qwen3-Next and performance update (#11969)
|
2025-10-30 21:52:16 +08:00 |
|
Jonah Bernard
|
62eff37ba1
|
Refactor Triton-kernel MoE runner integration (#11795)
|
2025-10-23 18:47:28 -07:00 |
|
Cheng Wan
|
bfc3b3f786
|
[9/N] MoE Refactor: cleanup dispatcher interfaces (#11847)
|
2025-10-20 10:11:46 -07:00 |
|
fzyzcjy
|
059c13de5c
|
Fix trtllm_moe wrong correction bias (#10440)
|
2025-09-15 01:02:05 -07:00 |
|
Cheng Wan
|
4844fac91d
|
Refactor TopK to ensure readability and extensibility (#9338)
|
2025-09-14 19:16:25 -07:00 |
|
Even Zhou
|
5b64f006ec
|
[Feature] Support DeepEP normal & Redundant Experts on NPU (#9881)
|
2025-09-10 20:35:26 -07:00 |
|
Guoyuan Lin
|
5e194b2143
|
[Model] Support Meituan LongCat-Flash && LongCat-Flash-MTP (#9824)
|
2025-08-30 23:29:21 -07:00 |
|
 chenxu140andEven Zhou
|
74dd4249ac
|
[Feature] Support NPUGraph for DeepSeek on Ascend NPU (#9355)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2025-08-28 16:06:24 -07:00 |
|
 ZhengdQinandEven Zhou
|
f92b729d52
|
[new feat] ascend backend support fia fusion kernel (#8328)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2025-08-25 23:13:08 -07:00 |
|
Kaixi Hou
|
18da2c96ec
|
[NVIDIA] Fix trtllm fp4 moe backend when used in MTP (#9384)
|
2025-08-21 00:54:01 -07:00 |
|
Cao E
|
c674bf9c6b
|
Fix biased_grouped_topk_cpu (#9420)
|
2025-08-20 19:18:48 -07:00 |
|
Trevor Morris
|
a91e90d9a3
|
[2/2] Fuse routed scaling factor into select_experts (#8690)
|
2025-08-20 15:10:16 -07:00 |
|
Jiaqi Gu
|
3c2c9f6c9e
|
[Bug] Fix input arguments of flashinfer_trtllm_moe (#9317)
|
2025-08-18 18:03:19 -07:00 |
|
Trevor Morris
|
eff4eb3fdd
|
Add fp4 quantize before all-gather for Flashinfer cutlass MoE DP (max throughput) (#7667)
|
2025-08-15 22:08:11 -07:00 |
|
Cheng Wan
|
295895120d
|
[6/N] MoE Refactor: Cleanup MoE-related configs (#8849)
|
2025-08-14 21:14:53 -07:00 |
|
Yineng Zhang
|
dd949ace23
|
Revert "[1/2][resubmit] sgl-kernel: Fuse routed scaling factor into m… (#9035)
|
2025-08-10 17:34:54 -07:00 |
|
Lianmin Zheng
|
ef48d5547e
|
Fix CI (#9013)
|
2025-08-09 16:00:10 -07:00 |
|
 
|
137e75daa1
|
[Feature] Optimize DeepSeek's DeepEP on Ascend NPU (#8355)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Hexq0210 <hexq0809521@gmail.com>
|
2025-08-09 01:35:00 -07:00 |
|
Trevor Morris
|
591c232f7c
|
[1/2][resubmit] sgl-kernel: Fuse routed scaling factor into moe_fused_gate (select_experts) (#8770)
|
2025-08-08 17:55:06 -07:00 |
|
 Even Zhouandronnie_zheng
|
fee0ab0fba
|
[CI] Ascend NPU CI enhancement (#8294)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2025-08-03 22:16:38 -07:00 |
|
Cheng Wan
|
b102353f8f
|
[MoE] Enable renormalize=False in Triton kernels (#8735)
|
2025-08-03 17:03:04 -07:00 |
|
Ke Bao
|
0242bb9c74
|
Fix triton kernels topk with keyword arguments (#8732)
|
2025-08-03 10:45:15 -07:00 |
|
erictanjn
|
a9dd3ec3e9
|
fix:reorder topk experts to ensure shared expert replaces minimal score (#8125)
|
2025-07-28 20:36:46 +08:00 |
|
  
|
b3eac168e7
|
Support triton kernels v3.4.0 for fused_moe (#8258)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Cheng Wan <cwan@x.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-07-27 02:28:49 -07:00 |
|
Lianmin Zheng
|
ed2e313eb6
|
Clean up server_args, triton cache manager (#8332)
|
2025-07-25 14:14:51 -07:00 |
|
Xiaoyu Zhang
|
9045cc1eb8
|
[torch.compile bug] avoid biased_grouped_topk_impl func repeatedly triggering torch.compile in forward pass (#8353)
|
2025-07-25 21:17:47 +08:00 |
|
Ke Bao
|
465968b2e3
|
Fix dtype error in CI (#8197)
|
2025-07-21 00:27:55 +08:00 |
|
Atream
|
a589a07167
|
fix moe gate dtype, fix tbo, fix fake dispatch (#7825)
|
2025-07-19 22:13:46 -07:00 |
|
Cheng Wan
|
15ad6c9086
|
[1/N] MoE Refactor: refactor select_experts (#7966)
|
2025-07-19 00:51:15 -07:00 |
|
jianan-gu
|
48c1fa7bb6
|
[CPU][Llama4] Fix Llama4 MoE inputs with "apply_router_weight_on_input" (#7889)
|
2025-07-17 21:43:25 -07:00 |
|
Cheng Wan
|
49b8777460
|
Refactor: move all quantization-related code to srt/layer/quantization (#7989)
|
2025-07-17 00:47:07 -07:00 |
|
 
|
766392c6bd
|
[feature]Ascend quantization support (#7791)
Co-authored-by: ichernob <ichernobnn@gmail.com>
Co-authored-by: liupeng <liupeng374@huawei.com>
|
2025-07-10 09:17:37 -07:00 |
|
jianan-gu
|
d389bedf72
|
[CPU][Qwen3 MoE] Enable fused_topk CPU fusion and enhance FP8 TP padding (#7838)
|
2025-07-09 02:04:21 -07:00 |
|
Ke Bao
|
8b1942c6cc
|
Remove type conversion and fix id map in topk (#7759)
|
2025-07-03 18:13:32 -07:00 |
|
 
|
489934be0a
|
fuse renormal into moe topk softmax kernel python code (#7751)
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: zhyncs <me@zhyncs.com>
|
2025-07-03 16:22:14 -07:00 |
|
 
|
1e0e549766
|
Ascend attention backend(PA&MLA) (#7722)
Co-authored-by: Maksim <makcum888e@mail.ru>
Co-authored-by: VDV1985 <vladdv85@mail.ru>
|
2025-07-03 09:23:19 -07:00 |
|
fzyzcjy
|
0c9c6c75a8
|
Move files related to EPLB (#7580)
|
2025-06-29 15:39:38 -07:00 |
|
valarLip
|
e984d5073b
|
enable aiter_biased_grouped_topk kernel (#7423)
|
2025-06-24 02:09:42 -07:00 |
|
  
|
094c116f7d
|
Update python API of activation, topk, norm and rope and remove vllm dependency (#6614)
Co-authored-by: Wu, Chunyuan <chunyuan.wu@intel.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
Co-authored-by: sdp <sdp@gnr799219.jf.intel.com>
|
2025-06-17 22:11:50 -07:00 |
|
fzyzcjy
|
da47621ccc
|
Minor speedup topk postprocessing (#7058)
|
2025-06-13 00:50:18 -07:00 |
|
fzyzcjy
|
2f715f51cc
|
Minor compile fused topk (#6944)
|
2025-06-07 01:40:38 -07:00 |
|
fzyzcjy
|
5aff1e9392
|
Fix Qwen3MoE missing token padding optimization (#6820)
|
2025-06-05 00:04:59 -07:00 |
|
Cheng Wan
|
81964328b7
|
Set num_fused_shared_experts as num_shared_experts when shared_experts fusion is not disabled (#6736)
|
2025-06-04 15:53:22 -07:00 |
|
Cheng Wan
|
8a5480528d
|
[Refactor] Rename n_share_experts_fusion as num_fused_shared_experts (#6735)
|
2025-06-03 17:48:24 -07:00 |
|