Commit Graph

11116 Commits

Author SHA1 Message Date
fzyzcjy
21af8e73ad Super tiny add comments to SGLANG_DEEPEP_NUM_MAX_DISPATCH_TOKENS_PER_RANK (#14048) 2025-11-27 22:16:43 +08:00
Baizhou Zhang
7ab548ef64 [2/2] Refactor DeepGeem requant for FP8 FusedMoE on Blackwell (#13960) 2025-11-27 09:00:26 -05:00
Liangsheng Yin
bab033b970 Adjust max-parallel for CUDA CI (#14057) 2025-11-27 20:48:37 +08:00
Yixin Dong
6350042696 feat: Naive support Spec V2 + Constrained Decoding (#13425)
Signed-off-by: Ubospica <ubospica@gmail.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2025-11-27 20:31:46 +08:00
fzyzcjy
25758647b1 Support sanity checking weight consistency especially for RL (#13854) 2025-11-27 20:25:12 +08:00
fzyzcjy
2bc8ee8b74 Tiny support 3D tensors in inverse_transform_scale_ue8m0 (#14002) 2025-11-27 20:20:45 +08:00
Jimmy
ab843ced31 [Feat]Add scheduler recv skipper weights to environment configuration (#13855) 2025-11-27 18:16:11 +08:00
Mick
6edffc6391 [diffusion] perf: improve black-forest-labs/FLUX.2-dev (#14040) 2025-11-27 14:49:52 +08:00
gaopengff
077ca70ee4 [Intel XPU]Add xpu support for get_device_memory_capacity (#13895) 2025-11-26 20:55:52 -08:00
Qiaolin Yu
7cb04dc0e5 Use trtllm mha decode kernel for target_verify in speculative decoding (#13976)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2025-11-26 20:40:34 -08:00
sunxxuns
5443db8759 fix: Fix AMD CI failures with HIP layernorm and PyPI connectivity (#13814)
Co-authored-by: root <root@mi300x8-005.atl1.do.cpe.ice.amd.com>
2025-11-27 11:30:37 +08:00
Stefan He
9f340ab1fb [Piecewise] support disable decode cuda graph when enable piecewise cuda graph (#13965) 2025-11-26 18:35:59 -08:00
Stefan He
70c6f95107 Add CODEOWNERS entry for batch_invariant_ops (#14026) 2025-11-26 18:35:37 -08:00
alisonshao
d941a3befa Fix nightly test failure: NSA indexer dtype (#14017) 2025-11-26 21:27:51 -05:00
Cheng Wan
b12c9e5c0a Fix installation for nvidia-nvshmem-cu12 (#14033) 2025-11-26 18:27:12 -08:00
Simo Lin
e9e90460ed [model-gateway] allow refill rate to be zero (#14030) 2025-11-26 17:15:27 -08:00
alisonshao
6330d6641b Fix flashinfer cutlass MoE output shape for non-FP4-packed inputs (#14028) 2025-11-26 18:09:02 -07:00
Sam
91e8dc371a [Feat][NVFP4] Enable NVFP4 MoE for Qwen series models (eg. Qwen3-Next) #13761 (#13761)
Co-authored-by: Kaixi Hou <kaixih@nvidia.com>
2025-11-26 17:53:45 -07:00
Lianmin Zheng
231df4b0d4 Cleanup server args (#14027) 2025-11-26 16:32:41 -08:00
alisonshao
b087ef8b4b Nightly test job filter (#14025) 2025-11-26 16:04:24 -08:00
ShawnY112358
5155016b56 [feat] update bucketed weights from distributed (#13824)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
2025-11-26 15:30:45 -08:00
Netanel Haber
082b54c689 Support nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16 (and nvidia/C-RADIOv2-H) (#12277) 2025-11-26 16:28:52 -07:00
alisonshao
a8ef4d1804 Add nightly test support to unified run_suite.py (#13941) 2025-11-26 15:16:12 -08:00
mrhaoxx
15ff6982b9 Support KTransformers for Qwen3-VL moe (#13983) 2025-11-26 14:59:15 -08:00
alisonshao
9adef42c60 Fix nightly test failure: CPP radix cache init (#14018) 2025-11-26 14:58:26 -08:00
Tianhao Zhou
685b9d82bd fix: cuda graph issue while running longcat_flash (#14007) 2025-11-26 14:57:22 -08:00
alisonshao
a223402ffb Add adapter_model.safetensors to corruption validation for LoRA (#14022) 2025-11-26 14:52:46 -08:00
Xinyue Zhang
44d0a848da [model-gateway] Fix flaky test_circuit_breaker_half_open_failure_reopens (#14019) 2025-11-26 14:52:10 -08:00
alisonshao
5b7da0f58e Temporarily disable test_update_weights_from_disk.py in CI (#14021) 2025-11-26 13:56:28 -08:00
Douglas Yang
697a77bf73 Add stress test workflow (#13937)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2025-11-26 13:24:39 -08:00
Lianmin Zheng
0a186924ba [Code sync] Fix registration of some ops in grok & Fix oss sync scripts (#13990)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
2025-11-26 13:11:52 -08:00
Stefan He
b6312e62ea Update CODEOWNERS for layer and executor files (#14020) 2025-11-26 12:59:04 -08:00
Kangyan-Zhou
779cbc6e4b Fix Nvidia nightly test trigger params when it is triggered by parent workflow (#13966) 2025-11-26 12:01:48 -08:00
Wenyi Xu
67c8c86722 [model-gateway][doc] Update transport terminology to protocol in README.md (#13872)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-26 11:44:59 -08:00
Simo Lin
69a03bc3f7 [ci] allow manual label to trigger ci in rust, change ci order (#14016) 2025-11-26 11:38:50 -08:00
Chang Su
66f242b98f [model gateway][grpc] Add tojson filter to override minijinja's tojson (#14013)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-26 11:05:01 -08:00
Baizhou Zhang
8a9b8b8457 Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" (#14015) 2025-11-26 10:45:23 -08:00
Simo Lin
7e964b5198 [ci] mark skip as success instead of failure (#14014) 2025-11-26 10:39:38 -08:00
Simo Lin
a0d9f6cdd8 [model-gateway] fix xpu ci (#14012) 2025-11-26 10:33:59 -08:00
Netanel Haber
8308cd3632 Support internvl on Blackwell (which doesn't support fa3): add SingletonCache support to Vision{Sdpa|Triton|Ascend}Attention (#13151) 2025-11-26 10:31:45 -08:00
Da Chen
e0e8a99630 fix: correct usage of minimax-m2 deepep moe forward (#13892)
Co-authored-by: Dash <dash@minimaxi.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
2025-11-26 10:01:23 -08:00
Mick
5102d00901 [diffusion] model: support black-forest-labs/FLUX.2-dev (#14000) 2025-11-27 01:49:40 +08:00
Liangsheng Yin
0dd759e087 Put pr-gate after check-changes (#14009) 2025-11-27 01:40:16 +08:00
Wenyi Xu
5e70880e64 [model-gateway] Add PostgreSQL support to binding (#13766)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-11-26 09:11:13 -08:00
Vedant V Jhaveri
9dab534b35 fix spec dec request level metrics (#13754)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2025-11-27 01:09:21 +08:00
Michelle Wu
262c3c1fde [Ascend] Support enable-mixed-chunk in non-MLA scenarios (#12491)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2025-11-27 00:00:21 +08:00
Liangsheng Yin
6c190cbda0 Rename: --hooks to --forward-hooks (#13994) 2025-11-26 22:26:28 +08:00
Liangsheng Yin
eff6a07c8f Fix get_load API (#13991) 2025-11-26 22:07:51 +08:00
Shangming Cai
b704b0a922 Optimize uneven PP layer distribution logic to improve PP performance (#13977) 2025-11-26 22:05:03 +08:00
vipwangerxiao
15729dbc8e Use dynamically maintained num_waiting_tokens in get_load() (#13203)
Signed-off-by: Peng Wang <peng_wang@linux.alibaba.com>
Co-authored-by: Peng Wang <peng_wang@linux.alibaba.com>
2025-11-26 20:34:45 +08:00