Commit Graph

7258 Commits

Author SHA1 Message Date
gaopengff
077ca70ee4 [Intel XPU]Add xpu support for get_device_memory_capacity (#13895) 2025-11-26 20:55:52 -08:00
Qiaolin Yu
7cb04dc0e5 Use trtllm mha decode kernel for target_verify in speculative decoding (#13976)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2025-11-26 20:40:34 -08:00
sunxxuns
5443db8759 fix: Fix AMD CI failures with HIP layernorm and PyPI connectivity (#13814)
Co-authored-by: root <root@mi300x8-005.atl1.do.cpe.ice.amd.com>
2025-11-27 11:30:37 +08:00
Stefan He
9f340ab1fb [Piecewise] support disable decode cuda graph when enable piecewise cuda graph (#13965) 2025-11-26 18:35:59 -08:00
Stefan He
70c6f95107 Add CODEOWNERS entry for batch_invariant_ops (#14026) 2025-11-26 18:35:37 -08:00
alisonshao
d941a3befa Fix nightly test failure: NSA indexer dtype (#14017) 2025-11-26 21:27:51 -05:00
Cheng Wan
b12c9e5c0a Fix installation for nvidia-nvshmem-cu12 (#14033) 2025-11-26 18:27:12 -08:00
Simo Lin
e9e90460ed [model-gateway] allow refill rate to be zero (#14030) 2025-11-26 17:15:27 -08:00
alisonshao
6330d6641b Fix flashinfer cutlass MoE output shape for non-FP4-packed inputs (#14028) 2025-11-26 18:09:02 -07:00
Sam
91e8dc371a [Feat][NVFP4] Enable NVFP4 MoE for Qwen series models (eg. Qwen3-Next) #13761 (#13761)
Co-authored-by: Kaixi Hou <kaixih@nvidia.com>
2025-11-26 17:53:45 -07:00
Lianmin Zheng
231df4b0d4 Cleanup server args (#14027) 2025-11-26 16:32:41 -08:00
alisonshao
b087ef8b4b Nightly test job filter (#14025) 2025-11-26 16:04:24 -08:00
ShawnY112358
5155016b56 [feat] update bucketed weights from distributed (#13824)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
2025-11-26 15:30:45 -08:00
Netanel Haber
082b54c689 Support nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16 (and nvidia/C-RADIOv2-H) (#12277) 2025-11-26 16:28:52 -07:00
alisonshao
a8ef4d1804 Add nightly test support to unified run_suite.py (#13941) 2025-11-26 15:16:12 -08:00
mrhaoxx
15ff6982b9 Support KTransformers for Qwen3-VL moe (#13983) 2025-11-26 14:59:15 -08:00
alisonshao
9adef42c60 Fix nightly test failure: CPP radix cache init (#14018) 2025-11-26 14:58:26 -08:00
Tianhao Zhou
685b9d82bd fix: cuda graph issue while running longcat_flash (#14007) 2025-11-26 14:57:22 -08:00
alisonshao
a223402ffb Add adapter_model.safetensors to corruption validation for LoRA (#14022) 2025-11-26 14:52:46 -08:00
Xinyue Zhang
44d0a848da [model-gateway] Fix flaky test_circuit_breaker_half_open_failure_reopens (#14019) 2025-11-26 14:52:10 -08:00
alisonshao
5b7da0f58e Temporarily disable test_update_weights_from_disk.py in CI (#14021) 2025-11-26 13:56:28 -08:00
Douglas Yang
697a77bf73 Add stress test workflow (#13937)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2025-11-26 13:24:39 -08:00
Lianmin Zheng
0a186924ba [Code sync] Fix registration of some ops in grok & Fix oss sync scripts (#13990)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
2025-11-26 13:11:52 -08:00
Stefan He
b6312e62ea Update CODEOWNERS for layer and executor files (#14020) 2025-11-26 12:59:04 -08:00
Kangyan-Zhou
779cbc6e4b Fix Nvidia nightly test trigger params when it is triggered by parent workflow (#13966) 2025-11-26 12:01:48 -08:00
Wenyi Xu
67c8c86722 [model-gateway][doc] Update transport terminology to protocol in README.md (#13872)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-26 11:44:59 -08:00
Simo Lin
69a03bc3f7 [ci] allow manual label to trigger ci in rust, change ci order (#14016) 2025-11-26 11:38:50 -08:00
Chang Su
66f242b98f [model gateway][grpc] Add tojson filter to override minijinja's tojson (#14013)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-26 11:05:01 -08:00
Baizhou Zhang
8a9b8b8457 Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" (#14015) 2025-11-26 10:45:23 -08:00
Simo Lin
7e964b5198 [ci] mark skip as success instead of failure (#14014) 2025-11-26 10:39:38 -08:00
Simo Lin
a0d9f6cdd8 [model-gateway] fix xpu ci (#14012) 2025-11-26 10:33:59 -08:00
Netanel Haber
8308cd3632 Support internvl on Blackwell (which doesn't support fa3): add SingletonCache support to Vision{Sdpa|Triton|Ascend}Attention (#13151) 2025-11-26 10:31:45 -08:00
Da Chen
e0e8a99630 fix: correct usage of minimax-m2 deepep moe forward (#13892)
Co-authored-by: Dash <dash@minimaxi.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
2025-11-26 10:01:23 -08:00
Mick
5102d00901 [diffusion] model: support black-forest-labs/FLUX.2-dev (#14000) 2025-11-27 01:49:40 +08:00
Liangsheng Yin
0dd759e087 Put pr-gate after check-changes (#14009) 2025-11-27 01:40:16 +08:00
Wenyi Xu
5e70880e64 [model-gateway] Add PostgreSQL support to binding (#13766)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-11-26 09:11:13 -08:00
Vedant V Jhaveri
9dab534b35 fix spec dec request level metrics (#13754)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2025-11-27 01:09:21 +08:00
Michelle Wu
262c3c1fde [Ascend] Support enable-mixed-chunk in non-MLA scenarios (#12491)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2025-11-27 00:00:21 +08:00
Liangsheng Yin
6c190cbda0 Rename: --hooks to --forward-hooks (#13994) 2025-11-26 22:26:28 +08:00
Liangsheng Yin
eff6a07c8f Fix get_load API (#13991) 2025-11-26 22:07:51 +08:00
Shangming Cai
b704b0a922 Optimize uneven PP layer distribution logic to improve PP performance (#13977) 2025-11-26 22:05:03 +08:00
vipwangerxiao
15729dbc8e Use dynamically maintained num_waiting_tokens in get_load() (#13203)
Signed-off-by: Peng Wang <peng_wang@linux.alibaba.com>
Co-authored-by: Peng Wang <peng_wang@linux.alibaba.com>
2025-11-26 20:34:45 +08:00
Zehuan Li
21b0582d4b [feature] Initial block diffusion language model support (#12588)
Co-authored-by: Tiwei Bie <tiwei.btw@antgroup.com>
2025-11-26 17:57:54 +08:00
Yuhao Yang
5795da5e83 [diffusion] fix: fix the issue where the qwen-edit & wan model produces incorrect output during sequence parallelism (#13922) 2025-11-26 15:58:23 +08:00
StonyPort
540d6fee20 Support piecewise CUDA graph for embedding models (#13852)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
2025-11-26 15:29:46 +08:00
Thomas Wang
5a8adca900 Turn off PREBUILD aiter in MI355 (#13963) 2025-11-25 22:30:39 -08:00
ShawnY112358
007c3e234c [feat] support in-flight weight update (#10071)
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
2025-11-25 22:03:13 -08:00
fzyzcjy
7130ad3a29 Fix SGLANG_ENABLE_HEALTH_ENDPOINT_GENERATION not working (#13961) 2025-11-25 21:58:55 -08:00
Zaili Wang
35a4c21a8a update CI permission list (#13962) 2025-11-25 20:16:44 -08:00
alisonshao
846ba3c62c Fix nightly test failures: NSA indexer dtype and CPP radix cache init (#13958) 2025-11-26 11:52:08 +08:00