Nicolas Castet
|
99df920cdb
|
Register tensors with symmetric memory for qwen (#18643)
|
2026-02-20 09:32:32 +08:00 |
|
Ke Bao
|
fb683be6eb
|
Use attn tp group in embedding for more models (#17570)
|
2026-01-24 13:37:44 +08:00 |
|
 Shifang XuandShu Wang
|
d27f16f38a
|
Fix EPLB + FP4 Quantization Compatibility Issue (#13715)
Co-authored-by: Shu Wang <shuw@nvidia.com>
|
2026-01-10 13:38:19 +08:00 |
|
 Nan JiangandXinyuan Tong
|
7254986342
|
[VLM] feat: true on policy for vlm + fsdp (#14636)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-01-01 16:54:39 -08:00 |
|
 
|
bed301a5ac
|
[Feature] Enable return routed experts (#12162)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-21 15:16:43 +08:00 |
|
 Yuhao YaoandCheng Wan
|
e9e7f15eb5
|
[bugfix] fix TBO crashes when attn_tp_size > 1 (#13730)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-12-11 18:18:40 -08:00 |
|
b8zhong
|
6d5d76ad97
|
remove unecessary dual stream token threshold from the rest of models (qwen moe, kimi linear, etc.) (#14337)
|
2025-12-06 19:57:26 -08:00 |
|
 jianan-guandZheng, Beilei <beilei.zheng@intel.com>
|
70d2587324
|
[CPU] Optimize small oc GEMM for Qwen3-next on CPU (#12446)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
|
2025-12-04 00:38:47 -08:00 |
|
Lianmin Zheng
|
ca52ed425f
|
Clean up imports and move files (#14317)
|
2025-12-02 16:31:54 -08:00 |
|
Sam
|
e7e89349c9
|
Enable Flashinfer TRTLLM-GEN-MoE FP8 blockwise kernel for Qwen3-Next on Blackwell (#12543)
|
2025-11-13 19:44:44 +08:00 |
|
Rain H
|
750940ae36
|
Eagle3 DP attention for Qwen3 MoE (#12002)
|
2025-10-29 20:25:17 +08:00 |
|
Xiaoyu Zhang
|
8374a96e49
|
piecewise cuda graph support qwen3-moe (#11845)
|
2025-10-21 10:55:49 +08:00 |
|
Cheng Wan
|
bfc3b3f786
|
[9/N] MoE Refactor: cleanup dispatcher interfaces (#11847)
|
2025-10-20 10:11:46 -07:00 |
|
 MickandXinyuan Tong
|
86b04d25b3
|
model: qwen3-omni (thinker-only) (#10911)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-10-16 13:20:38 -07:00 |
|
Liangsheng Yin
|
516738b096
|
Depreate global_server_args_dict (#11528)
|
2025-10-13 19:34:43 +08:00 |
|
Cheng Wan
|
1bdd010291
|
Revert "Deprecate global_server_args_dict" (#11520)
|
2025-10-12 17:40:40 -07:00 |
|
Liangsheng Yin
|
1083e7e3df
|
Deprecate global_server_args_dict (#11331)
|
2025-10-13 01:20:47 +08:00 |
|
Yi Zhang
|
1344ebc833
|
support qwen3-next-fp8 deepep (#10622)
|
2025-09-18 11:36:22 -07:00 |
|
Cheng Wan
|
4844fac91d
|
Refactor TopK to ensure readability and extensibility (#9338)
|
2025-09-14 19:16:25 -07:00 |
|
Yi Zhang
|
27778010fc
|
fix dual stream bug (#10352)
|
2025-09-11 20:53:42 -07:00 |
|
Yi Zhang
|
9e2f7252db
|
add dual stream for qwen2_moe (#10252)
|
2025-09-10 12:49:43 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
ec15c8360e
|
Optimize Qwen3-moe model by using flashinfer fused allreduce (#9973)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-09-04 20:48:53 +08:00 |
|
KerwinKai
|
87a0f7d2c2
|
[feat] Support EAGLE3 for Qwen2 (#9216)
|
2025-08-29 12:59:51 -07:00 |
|
Cheng Wan
|
295895120d
|
[6/N] MoE Refactor: Cleanup MoE-related configs (#8849)
|
2025-08-14 21:14:53 -07:00 |
|
wxzhoucs
|
4c22897a66
|
Feature: support qwen and llama4 reducescatter for dp attention padding (#9101)
|
2025-08-13 21:10:29 -07:00 |
|
Cheng Wan
|
b87aacb5c5
|
[DP Attention] Refactor: adding some utility functions (#9136)
|
2025-08-13 21:08:06 -07:00 |
|
PGFLMG
|
b7cd743038
|
[Feat] QWen-1M context support[2/2]: Update block sparse attention backend (#5949)
|
2025-08-06 23:49:36 -07:00 |
|
Cheng Wan
|
6c88f6c8d9
|
[5/N] MoE Refactor: Update MoE parallelism arguments (#8658)
|
2025-08-01 01:20:03 -07:00 |
|
Kaixi Hou
|
85486b6f6f
|
[NVIDIA] Add Flashinfer MoE blockscale fp8 backend (#8036)
|
2025-07-27 00:34:41 -07:00 |
|
Cheng Wan
|
c0fb25e949
|
DP Enhancement (#8280)
|
2025-07-24 21:36:21 -07:00 |
|
Cheng Wan
|
15ad6c9086
|
[1/N] MoE Refactor: refactor select_experts (#7966)
|
2025-07-19 00:51:15 -07:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) Xiaoze Fanandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
570d33437b
|
[Feature] Layer-wise Prefill (#7634)
Signed-off-by: jason-fxz <jason341132@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-07-17 01:57:46 +08:00 |
|
 
|
1964c325de
|
[feat] Support EAGLE3 for Qwen (#7745)
Co-authored-by: 纬杭 <ximing.wxm@antgroup.com>
Co-authored-by: zyksir <zyksir@outlook.com>
|
2025-07-04 19:50:28 -07:00 |
|
yilian49
|
c01a1df588
|
[Bug] add flashinfer bool check for fusedmoe in Qwen moe models (#7723)
|
2025-07-03 11:32:11 -07:00 |
|
 Yi Zhangandispobock
|
264dc6e744
|
[optimize] add two stream norm for qwen3 (#7740)
Co-authored-by: ispobock <ispobaoke@gmail.com>
|
2025-07-03 09:59:17 -07:00 |
|
fzyzcjy
|
0c9c6c75a8
|
Move files related to EPLB (#7580)
|
2025-06-29 15:39:38 -07:00 |
|
Yineng Zhang
|
fa6723f08f
|
Revert "fix communicator for non-dp lm head (#6662)" (#6677)
|
2025-05-27 12:22:59 -07:00 |
|
Cheng Wan
|
a3d7f4b673
|
fix communicator for non-dp lm head (#6662)
|
2025-05-27 02:31:12 -07:00 |
|
Yi Zhang
|
b18416fbf8
|
Fix qwen3 tbo/dp-lm-head (#6652)
|
2025-05-27 00:38:27 -07:00 |
|
fzyzcjy
|
32cd707002
|
Support TP in attention for two batch overlap (#6634)
|
2025-05-26 20:28:12 -07:00 |
|
Yi Zhang
|
f9bab3d591
|
qwen3moe support two batch overlap (#6598)
|
2025-05-25 23:08:16 -07:00 |
|
Yi Zhang
|
65f091310c
|
refactor qwen moe code, use communicator to support tp+dp (#6581)
|
2025-05-25 23:01:10 -07:00 |
|
lukec
|
fc0e3b9174
|
Support qwen3 deepep (#6120)
|
2025-05-22 11:04:45 -07:00 |
|
fzyzcjy
|
f0653886a5
|
Expert distribution recording without overhead for EPLB (#4957)
|
2025-05-19 20:07:43 -07:00 |
|
libra
|
11553c1a37
|
Add pipeline parallelism for Qwen2 and Qwen3 Model (#6250)
|
2025-05-18 00:42:55 -07:00 |
|
 
|
4bd2952a37
|
feat: add dp attention support for Qwen 2/3 MoE models, fixes #6088 (#6121)
Co-authored-by: King.Zevin <zevin@mail.ustc.edu.cn>
Co-authored-by: Yi Zhang <1109276519@qq.com>
|
2025-05-16 14:44:10 -07:00 |
|
 laixinandsleepcoo
|
e330f2b86c
|
[qwen3] support qwen3 ep moe (#5917)
Co-authored-by: sleepcoo <sleepcoo@gmail.com>
|
2025-04-30 09:15:21 -07:00 |
|
yhyang201
|
4db463b1ad
|
[Model] Adding Qwen3 and Qwen3MoE (#4693)
|
2025-04-18 09:51:29 -07:00 |
|
Michael Feil
|
1effba4c70
|
Configuration qwen2_moe.py - qkv_bias now in transformers (#5512)
|
2025-04-17 21:23:22 -07:00 |
|
 Yun Daiandqingquansong
|
2695ab0537
|
Fix loading KV quantization scale; Enable modelopt kv cache (#4686)
Co-authored-by: qingquansong <ustcsqq@gmail.com>
|
2025-04-08 09:11:35 -07:00 |
|