tql.99
|
3f2e315f6e
|
optimize: reduce shulffle and quantization overhead in cutlass_moe sm90 (#8962)
Co-authored-by: 戚余航 <qiyuhang@bytedance.com>
|
2025-08-09 00:29:12 -07:00 |
|
Lianmin Zheng
|
706bd69cc5
|
Clean up server_args.py to have a dedicated function for model specific adjustments (#8983)
|
2025-08-08 19:56:50 -07:00 |
|
Cheng Wan
|
1d24db8348
|
Expert Parallelism for GPT-OSS (#8944)
|
2025-08-08 00:46:42 -07:00 |
|
Xiaoyu Zhang
|
0d1e27a0c5
|
Better optimization log for gpt-oss model (#8953)
|
2025-08-08 00:11:48 -07:00 |
|
Stefan He
|
aaf0ad8cdf
|
remove vllm fp8quant from fp8.py (#8937)
|
2025-08-07 15:50:52 -07:00 |
|
Xiaoyu Zhang
|
3ae33fcd0a
|
Fix hopper launch gpt-oss model illegal memory (#8908)
|
2025-08-07 10:02:40 -07:00 |
|
Xiaoyu Zhang
|
47824c1488
|
[Perf] Auto enable best flashinfer mxfp4 kernel in b200 (#8898)
|
2025-08-07 01:08:41 -07:00 |
|
Xiaoyu Zhang
|
4373df5525
|
add flashinfer mxfp4 (#8847)
|
2025-08-06 16:23:41 -07:00 |
|
Shu Wang
|
288ae41f7a
|
[NVIDIA] Fix num_experts in modelopt_quant (#8811)
|
2025-08-06 14:35:07 -07:00 |
|
Ying Sheng
|
168033d5fb
|
Support mxfp4 for GPT-OSS (#8843)
Co-authored-by: Co-author fzyzcjy <ch271828n@outlook.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Co-authored-by: zhuofan1123 <zhuofanl@nvidia.com>
Co-authored-by: liz-badada <jinyanc@nvidia.com>
Co-authored-by: xutizhou <xutingz@nvidia.com>
Co-authored-by: linhu-nv <linhu@nvidia.com>
|
2025-08-06 00:05:25 -07:00 |
|
Ying Sheng
|
c1d2061f97
|
Add initial support for gpt-oss (#8824)
|
2025-08-05 13:42:01 -07:00 |
|
Yineng Zhang
|
5e91fed1c5
|
Revert "[NVIDIA]Fix local_num_experts for EP (#8779)" (#8797)
|
2025-08-04 23:30:43 -07:00 |
|
Shu Wang
|
b01eeb80f8
|
[NVIDIA]Fix local_num_experts for EP (#8779)
|
2025-08-04 22:01:14 -07:00 |
|
kk
|
d4bf5a8524
|
Support OCP MXFP4 quantization on AMD GPUs (#8255)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
|
2025-08-04 18:14:52 -07:00 |
|
Trevor Morris
|
9bd4872a34
|
[bugfix] Fix typo in modelopt quant: 'FusedMoE' object has no attribute 'local_num_experts' (#8768)
|
2025-08-04 11:08:08 -07:00 |
|
azhurkevich
|
915140fd18
|
[NVIDIA] Add Low Latency NVFP4 decode kernels from Flashinfer (#8552)
Co-authored-by: Cheng Wan <cwan@x.ai>
|
2025-08-04 03:10:02 -07:00 |
|
Yuan Luo
|
3b87a9e8ae
|
Fix bug of refactoring TopKOutput in w4afp8 (#8745)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-08-03 20:05:02 -07:00 |
|
tql.99
|
e67276ecb3
|
feat: support cutlass_moe_fp8 kernel for fusedmoe in sm90 (#8678)
|
2025-08-03 10:47:15 -07:00 |
|
Cheng Wan
|
a437aa9987
|
[hotfix] fix mixtral with tensor-level compressed-tensor quantization (#8721)
|
2025-08-02 22:59:25 -07:00 |
|
fzyzcjy
|
403566bcca
|
Remove assertions about per group quant fp8 (#8717)
|
2025-08-02 17:08:40 -07:00 |
|
Trevor Morris
|
89caf7a3c6
|
[bugfix] Apply routed scaling factor to cutlass_fused_experts_fp8 (#8688)
|
2025-08-01 19:00:24 -07:00 |
|
Ke Bao
|
e252192679
|
Fix deepgemm masked grouped gemm jit compile (#8679)
|
2025-08-01 15:37:59 -07:00 |
|
Kaixi Hou
|
aa4c66b564
|
[NVIDIA] Enable Flashinfer MoE blockscale fp8 backend for TP MoE (#8450)
Co-authored-by: kushanam <42385577+kushanam@users.noreply.github.com>
|
2025-07-31 19:56:34 -07:00 |
|
Even Zhou
|
99795d61e6
|
[Bugfix] fix w8a8_int8 load issue (#8308)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2025-07-31 17:30:16 -07:00 |
|
Cheng Wan
|
32fa1e9cc2
|
[4/N] MoE Refactor: Unified Triton Kernel for FusedMoE and EPMoE (#8515)
|
2025-07-31 02:34:02 -07:00 |
|
fzyzcjy
|
7df2c0c2db
|
Reduce memory usage for fp4 moe (#8413)
|
2025-07-28 22:51:23 -07:00 |
|
fzyzcjy
|
b58c3c285e
|
Support ue8m0 for triton quant kernel (#7603)
|
2025-07-27 13:04:35 -07:00 |
|
Yuan Luo
|
b3eac168e7
|
Support triton kernels v3.4.0 for fused_moe (#8258)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Cheng Wan <cwan@x.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-07-27 02:28:49 -07:00 |
|
Elfie Guo
|
5c9c275bc8
|
Use FlashInfer FP4 gemm. (#8241)
|
2025-07-27 01:05:22 -07:00 |
|
Cheng Wan
|
bf0f448fe5
|
[2/N] MoE Refactor: Unify weight loader and quant methods (#8397)
|
2025-07-27 01:00:21 -07:00 |
|
Kaixi Hou
|
85486b6f6f
|
[NVIDIA] Add Flashinfer MoE blockscale fp8 backend (#8036)
|
2025-07-27 00:34:41 -07:00 |
|
Trevor Morris
|
58c468f404
|
Fix FP4 MoE accuracy from missing routed_scaling_factor (#8333)
|
2025-07-25 16:40:23 -07:00 |
|
Xiaoyu Zhang
|
a167fd0bcb
|
[code style] Clean dead triton kernel code in fused_moe and useless vllm_ops import (#8310)
|
2025-07-24 14:38:30 +08:00 |
|
Hubert Lu
|
e50109f2ed
|
[AMD] Remove vllm's scaled_fp8_quant and moe_sum when SGLANG_USE_AITER=1 (#7484)
|
2025-07-21 17:33:19 -07:00 |
|
JieXin Liang
|
7eebd44047
|
[fix] fix modelopt fp4 on b200 (#8195)
|
2025-07-20 17:39:57 -07:00 |
|
Cheng Wan
|
15ad6c9086
|
[1/N] MoE Refactor: refactor select_experts (#7966)
|
2025-07-19 00:51:15 -07:00 |
|
Haohui Mai
|
d918ab7985
|
Support NVFP4 quantized dense models on AMD CDNA2/CDNA3 GPUs (#7302)
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
|
2025-07-18 19:59:39 -07:00 |
|
Hubert Lu
|
7750b91ca8
|
[AMD] Add triton awq_dequantize kernel to support AWQ on ROCm (#7661)
|
2025-07-18 14:27:25 -07:00 |
|
Hongbo Xu
|
1f76fc8747
|
[3/n] chore: decouple AWQ implementation from vLLM dependency (#8113)
Co-authored-by: AniZpZ <zhuangsen.zp@antgroup.com>
|
2025-07-18 11:45:22 -07:00 |
|
Even Zhou
|
6737671c82
|
[Bugfix] Fix w8a8_int8 import error on NPU (#8147)
|
2025-07-18 11:34:55 -07:00 |
|
Enrique Shockwave
|
fd63b62eaa
|
fix compressed tensors WNA16 imports (#8142)
|
2025-07-18 11:34:14 -07:00 |
|
jianan-gu
|
7891bac16b
|
[Quantization][w8a8_int8] Fix weight loading issue for w8a8_int8 path with "ignore" layer list in quantization config (#7820)
|
2025-07-17 22:03:56 -07:00 |
|
jianan-gu
|
48c1fa7bb6
|
[CPU][Llama4] Fix Llama4 MoE inputs with "apply_router_weight_on_input" (#7889)
|
2025-07-17 21:43:25 -07:00 |
|
Cheng Wan
|
49b8777460
|
Refactor: move all quantization-related code to srt/layer/quantization (#7989)
|
2025-07-17 00:47:07 -07:00 |
|
Peng Zhang
|
c28ad1990d
|
[1/n] chore: decouple quantization implementation from vLLM dependency (#7992)
|
2025-07-16 15:56:26 -07:00 |
|
YanbingJiang
|
b188a89a5d
|
Fix CI xeon test with triton 3.3.1 (#8086)
|
2025-07-16 02:12:23 -07:00 |
|
Qi Yuhang
|
c268c11c71
|
[feat]Support fusion kernel for constructing quant input and scale factor for fp8_blockwise_scaled_grouped_mm (#8023)
|
2025-07-15 00:02:44 -07:00 |
|
mqhc2020
|
a562c8a35c
|
[Dockerfile] Multi-arch support for ROCm (#7902)
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2025-07-14 06:13:09 +00:00 |
|
ronnie_zheng
|
766392c6bd
|
[feature]Ascend quantization support (#7791)
Co-authored-by: ichernob <ichernobnn@gmail.com>
Co-authored-by: liupeng <liupeng374@huawei.com>
|
2025-07-10 09:17:37 -07:00 |
|
likesen-alibaba
|
4a0d19198b
|
Fix bug of deepseek-v3 under DP+EP mode with large batchsize/seqlen (#6449)
|
2025-07-10 01:19:56 -07:00 |
|