fzyzcjy
21af8e73ad
Super tiny add comments to SGLANG_DEEPEP_NUM_MAX_DISPATCH_TOKENS_PER_RANK ( #14048 )
2025-11-27 22:16:43 +08:00
Baizhou Zhang
7ab548ef64
[2/2] Refactor DeepGeem requant for FP8 FusedMoE on Blackwell ( #13960 )
2025-11-27 09:00:26 -05:00
Liangsheng Yin
bab033b970
Adjust max-parallel for CUDA CI ( #14057 )
2025-11-27 20:48:37 +08:00
Yixin Dong
6350042696
feat: Naive support Spec V2 + Constrained Decoding ( #13425 )
...
Signed-off-by: Ubospica <ubospica@gmail.com >
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com >
2025-11-27 20:31:46 +08:00
fzyzcjy
25758647b1
Support sanity checking weight consistency especially for RL ( #13854 )
2025-11-27 20:25:12 +08:00
fzyzcjy
2bc8ee8b74
Tiny support 3D tensors in inverse_transform_scale_ue8m0 ( #14002 )
2025-11-27 20:20:45 +08:00
Jimmy
ab843ced31
[Feat]Add scheduler recv skipper weights to environment configuration ( #13855 )
2025-11-27 18:16:11 +08:00
Mick
6edffc6391
[diffusion] perf: improve black-forest-labs/FLUX.2-dev ( #14040 )
2025-11-27 14:49:52 +08:00
gaopengff
077ca70ee4
[Intel XPU]Add xpu support for get_device_memory_capacity ( #13895 )
2025-11-26 20:55:52 -08:00
Qiaolin Yu
7cb04dc0e5
Use trtllm mha decode kernel for target_verify in speculative decoding ( #13976 )
...
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com >
2025-11-26 20:40:34 -08:00
sunxxuns
5443db8759
fix: Fix AMD CI failures with HIP layernorm and PyPI connectivity ( #13814 )
...
Co-authored-by: root <root@mi300x8-005.atl1.do.cpe.ice.amd.com >
2025-11-27 11:30:37 +08:00
Stefan He
9f340ab1fb
[Piecewise] support disable decode cuda graph when enable piecewise cuda graph ( #13965 )
2025-11-26 18:35:59 -08:00
Stefan He
70c6f95107
Add CODEOWNERS entry for batch_invariant_ops ( #14026 )
2025-11-26 18:35:37 -08:00
alisonshao
d941a3befa
Fix nightly test failure: NSA indexer dtype ( #14017 )
2025-11-26 21:27:51 -05:00
Cheng Wan
b12c9e5c0a
Fix installation for nvidia-nvshmem-cu12 ( #14033 )
2025-11-26 18:27:12 -08:00
Simo Lin
e9e90460ed
[model-gateway] allow refill rate to be zero ( #14030 )
2025-11-26 17:15:27 -08:00
alisonshao
6330d6641b
Fix flashinfer cutlass MoE output shape for non-FP4-packed inputs ( #14028 )
2025-11-26 18:09:02 -07:00
Sam
91e8dc371a
[Feat][NVFP4] Enable NVFP4 MoE for Qwen series models (eg. Qwen3-Next) #13761 ( #13761 )
...
Co-authored-by: Kaixi Hou <kaixih@nvidia.com >
2025-11-26 17:53:45 -07:00
Lianmin Zheng
231df4b0d4
Cleanup server args ( #14027 )
2025-11-26 16:32:41 -08:00
alisonshao
b087ef8b4b
Nightly test job filter ( #14025 )
2025-11-26 16:04:24 -08:00
ShawnY112358
5155016b56
[feat] update bucketed weights from distributed ( #13824 )
...
Co-authored-by: Stefan He <hebiaobuaa@gmail.com >
2025-11-26 15:30:45 -08:00
Netanel Haber
082b54c689
Support nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16 (and nvidia/C-RADIOv2-H) ( #12277 )
2025-11-26 16:28:52 -07:00
alisonshao
a8ef4d1804
Add nightly test support to unified run_suite.py ( #13941 )
2025-11-26 15:16:12 -08:00
mrhaoxx
15ff6982b9
Support KTransformers for Qwen3-VL moe ( #13983 )
2025-11-26 14:59:15 -08:00
alisonshao
9adef42c60
Fix nightly test failure: CPP radix cache init ( #14018 )
2025-11-26 14:58:26 -08:00
Tianhao Zhou
685b9d82bd
fix: cuda graph issue while running longcat_flash ( #14007 )
2025-11-26 14:57:22 -08:00
alisonshao
a223402ffb
Add adapter_model.safetensors to corruption validation for LoRA ( #14022 )
2025-11-26 14:52:46 -08:00
Xinyue Zhang
44d0a848da
[model-gateway] Fix flaky test_circuit_breaker_half_open_failure_reopens ( #14019 )
2025-11-26 14:52:10 -08:00
alisonshao
5b7da0f58e
Temporarily disable test_update_weights_from_disk.py in CI ( #14021 )
2025-11-26 13:56:28 -08:00
Douglas Yang
697a77bf73
Add stress test workflow ( #13937 )
...
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com >
2025-11-26 13:24:39 -08:00
Lianmin Zheng
0a186924ba
[Code sync] Fix registration of some ops in grok & Fix oss sync scripts ( #13990 )
...
Co-authored-by: Stefan He <hebiaobuaa@gmail.com >
2025-11-26 13:11:52 -08:00
Stefan He
b6312e62ea
Update CODEOWNERS for layer and executor files ( #14020 )
2025-11-26 12:59:04 -08:00
Kangyan-Zhou
779cbc6e4b
Fix Nvidia nightly test trigger params when it is triggered by parent workflow ( #13966 )
2025-11-26 12:01:48 -08:00
Wenyi Xu
67c8c86722
[model-gateway][doc] Update transport terminology to protocol in README.md ( #13872 )
...
Co-authored-by: Simo Lin <linsimo.mark@gmail.com >
2025-11-26 11:44:59 -08:00
Simo Lin
69a03bc3f7
[ci] allow manual label to trigger ci in rust, change ci order ( #14016 )
2025-11-26 11:38:50 -08:00
Chang Su
66f242b98f
[model gateway][grpc] Add tojson filter to override minijinja's tojson ( #14013 )
...
Co-authored-by: Simo Lin <linsimo.mark@gmail.com >
2025-11-26 11:05:01 -08:00
Baizhou Zhang
8a9b8b8457
Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" ( #14015 )
2025-11-26 10:45:23 -08:00
Simo Lin
7e964b5198
[ci] mark skip as success instead of failure ( #14014 )
2025-11-26 10:39:38 -08:00
Simo Lin
a0d9f6cdd8
[model-gateway] fix xpu ci ( #14012 )
2025-11-26 10:33:59 -08:00
Netanel Haber
8308cd3632
Support internvl on Blackwell (which doesn't support fa3): add SingletonCache support to Vision{Sdpa|Triton|Ascend}Attention ( #13151 )
2025-11-26 10:31:45 -08:00
Da Chen
e0e8a99630
fix: correct usage of minimax-m2 deepep moe forward ( #13892 )
...
Co-authored-by: Dash <dash@minimaxi.com >
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com >
2025-11-26 10:01:23 -08:00
Mick
5102d00901
[diffusion] model: support black-forest-labs/FLUX.2-dev ( #14000 )
2025-11-27 01:49:40 +08:00
Liangsheng Yin
0dd759e087
Put pr-gate after check-changes ( #14009 )
2025-11-27 01:40:16 +08:00
Wenyi Xu
5e70880e64
[model-gateway] Add PostgreSQL support to binding ( #13766 )
...
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com >
Co-authored-by: Chang Su <chang.s.su@oracle.com >
2025-11-26 09:11:13 -08:00
Vedant V Jhaveri
9dab534b35
fix spec dec request level metrics ( #13754 )
...
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com >
2025-11-27 01:09:21 +08:00
Michelle Wu
262c3c1fde
[Ascend] Support enable-mixed-chunk in non-MLA scenarios ( #12491 )
...
Co-authored-by: ronnie_zheng <zl19940307@163.com >
2025-11-27 00:00:21 +08:00
Liangsheng Yin
6c190cbda0
Rename: --hooks to --forward-hooks ( #13994 )
2025-11-26 22:26:28 +08:00
Liangsheng Yin
eff6a07c8f
Fix get_load API ( #13991 )
2025-11-26 22:07:51 +08:00
Shangming Cai
b704b0a922
Optimize uneven PP layer distribution logic to improve PP performance ( #13977 )
2025-11-26 22:05:03 +08:00
vipwangerxiao
15729dbc8e
Use dynamically maintained num_waiting_tokens in get_load() ( #13203 )
...
Signed-off-by: Peng Wang <peng_wang@linux.alibaba.com >
Co-authored-by: Peng Wang <peng_wang@linux.alibaba.com >
2025-11-26 20:34:45 +08:00