Commit Graph
4793 Commits
Author SHA1 Message Date
lif 5969be2f06 Apply fixture-kit mode to MMMUVLMMixin (#15615) 2025-12-28 17:22:05 +08:00
Liangsheng Yin bf90ea9c5b Unify spec v2's naming manner. (#15990) 2025-12-28 14:14:52 +08:00
TZHelloWorld 0294844f04 [fix]deepgemm precompile when warmup (#15891) 2025-12-28 13:47:39 +08:00
Alison ShaoandKangyan-Zhou 0e536600e8 Refactor: separate CI-specific weight validation into dedicated module (#15216)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2025-12-27 20:50:39 -08:00
Vladislav Nosivskoyandishandhanani d70c265533 SGLang Tracing: fix attribute errors (header extraction & bootstrap span closing) (#15693)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2025-12-27 22:44:31 -06:00
Baizhou Zhang 656f4d69a1 Refactor fp8 nextn layer for DeepSeek nvfp4 checkpoint (#15353) 2025-12-28 11:57:09 +08:00
DarkSharpness 8e43980ebb [Feature] JIT Fused QK norm + qk norm clean up (#15835) 2025-12-28 11:53:50 +08:00
fzyzcjy f55d608c89 Tiny fix cannot launch nvfp4 checkpoint with bf16 kv cache (#15986) 2025-12-28 11:22:20 +08:00
Ho-Ren (Jack) Chuang 349ce2dd19 Support kv8 (FP8) with torch_native attention backend (#12596)
Signed-off-by: Ho-Ren (Jack) Chuang <horenchuang@bytedance.com>
2025-12-28 10:48:30 +08:00
Lianmin Zheng 183b65190a Clean up logging (#15919) 2025-12-27 15:27:12 -08:00
Hudson Xing 5c393e8153 Fix temp_prefill_info assertion error in PP disaggregation mode (#15943) 2025-12-27 22:46:55 +08:00
Ratish P ca740a41f3 [model-gateway]: fix grpc embedding test (#15934) 2025-12-27 05:56:54 -05:00
jiaming1130andZhengdQin 60a230b1fd [NPU] Support w4a8 with activation clip (#14736)
Co-authored-by: ZhengdQin <46387172+ZhengdQin@users.noreply.github.com>
2025-12-27 16:19:46 +08:00
Lianmin Zheng a8380ded71 Add a test case for crash dump (#15905) 2025-12-26 19:39:28 -08:00
Zheng Wengang cd3289c7a4 [BugFix][VLM] Correct weight loading with tie_word_embeddings == False (#15398) 2025-12-26 17:30:50 -08:00
Cheng Wan 2ec57cefd9 hotfix: add type hints to scheduler mixins (#15916) 2025-12-26 17:08:11 -08:00
Cheng Wan 988b14ca0e refactor: add type hints to scheduler mixins (#15913) 2025-12-26 16:50:07 -08:00
Lianmin Zheng 93495dcac9 Revert "[VLM] Refactor load_mm_data to improve performance" (#15911) 2025-12-26 13:43:18 -08:00
Ratish P 886e038329 [model-gateway]: fix crash in embedding worker health check (#15910) 2025-12-26 12:57:13 -08:00
Muqi Li 01bd0d3e8b [Tool Call][DSV32] Streamline function call parameters (#14750)
Signed-off-by: Muqi Li <muqi1029@gmail.com>
2025-12-26 11:35:30 -08:00
Huaixin ChangandJunjie Mao cf34d0ab32 [Fix] assert error in log_prefill_stats (#15881)
Signed-off-by: Chang Huaixin (OpenAnolis) <changhuaixin@linux.alibaba.com>
Co-authored-by: Junjie Mao (Alibaba) <banxing.mjj@alibaba-inc.com>
2025-12-26 19:54:35 +08:00
fzyzcjy 5d421db883 Tiny log warn users when tracing is automatically disabled (#15889) 2025-12-26 19:00:37 +08:00
shuwenn fe3d47fc9d fix: warn once per env var key (#15846) 2025-12-26 18:41:15 +08:00
Yi Zhang ef92b4eb88 [BUGFIX] fix edge case for qwen3-next (#14209) 2025-12-26 17:14:37 +08:00
Liangsheng Yin e75657c839 [Bug] fix piggyback load report return None bug (#15870) 2025-12-26 14:48:45 +08:00
Ke Bao c28c536c91 Fix swa available memory check (#15867) 2025-12-26 13:19:01 +08:00
Yuan Luoandluoyuan.luo 086813ae8a [VLM] refactor: refactor load_mm_data to improve performance (#14644)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-12-26 13:17:53 +08:00
Hexq0210 3fd232ad96 [NPU] update Mixed chunk op to FIA (#15518) 2025-12-26 12:17:40 +08:00
Liangsheng Yin cb1812954a Introduce ModelRunnerKVCacheMixin to simplify the code. (#15821) 2025-12-26 11:17:33 +08:00
Kangyan-Zhou a91e072f33 Use allow auto truncate in the OpenAI API endpoint (#15369) 2025-12-25 18:45:57 -08:00
fzyzcjy 0271fc3456 Fix prefill num tokens metrics (#15858) 2025-12-26 10:26:42 +08:00
fzyzcjy 68bece8caa Super tiny move last_prefill_tokens to metrics mixin (#15857) 2025-12-26 10:26:10 +08:00
Ke Bao 7b7e357f61 Separate swa and local attention chunk cache eviction (#15820) 2025-12-26 09:34:22 +08:00
Ke Bao 2f66b0671b Fix chunk_kda_fwd missing argument (#15851) 2025-12-26 09:32:35 +08:00
Hudson Xing 9d878c1f3e Optimize FP8 MLA KV cache writes with Triton kernel (#15522) 2025-12-25 12:35:39 -08:00
Lianmin Zheng 8087ef126f Clean up the __init__ of TokenizerManager and DetokenizerManager (#15796) 2025-12-25 10:30:27 -08:00
Yuxuan Zhang f3ba711662 fix: change class name of GLM-ASR (#15772) 2025-12-26 00:15:16 +08:00
Yuwei An 5c243ba588 Custom All Reduce for Piecewise Cuda Graph (#15356)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
2025-12-25 23:55:15 +08:00
b6702d72cf feat: log request when e2e latency exceeds the specified value (#15759)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2025-12-25 23:12:35 +08:00
Yingchun Lai bb9e6cdf9e [MiMoV2Flash] fix: respect --swa-full-tokens-ratio arg (#15488) 2025-12-25 21:02:56 +08:00
roikoren755 de03b0cd30 [Nemotron 3 Nano] Add triton MoE configs (#15815)
Signed-off-by: Roi Koren <roik@nvidia.com>
2025-12-25 19:31:20 +08:00
Liangsheng Yin f4e835af2f Cleanup ModelRunner (#15802) 2025-12-25 18:13:30 +08:00
a89e85e739 [1/N][Sparse With Hicache]: Add Sparse Interface (#14741)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: MagicYang1573 <1328657938@qq.com>
2025-12-25 00:12:47 -08:00
Ke BaoandLiangsheng Yin cbf9f13493 Adjust server args for Mimo-v2-flash model (#15803)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2025-12-25 16:04:39 +08:00
Yi Zhang eb3da9c1dd fuse ssm state store into chunk_gated_delta_rule_fwd_h (#15409) 2025-12-25 16:03:15 +08:00
Yuan Luoandluoyuan.luo b9af8d2eb9 [VLM] Support apply qk norm in multi cuda streams (#15720)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-12-25 14:35:07 +08:00
Liangsheng Yin b311c43d13 Clarify None handling in sglang's environ (#15770) 2025-12-25 11:40:17 +08:00
Huaixin Chang 0c39730b18 DP: support piggyback server load report (#11469)
Signed-off-by: Chang Huaixin (OpenAnolis) <changhuaixin@linux.alibaba.com>
2025-12-25 11:35:05 +08:00
satyamk7054andSatyam Kumar 38dd4fbb66 Add overlap scheduling for embeddings code path (#14032)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
2025-12-24 18:24:18 -08:00
Lianmin Zhenggithub-actions[bot] <github-actions[bot]@users.noreply.github.com>Stefan He
92ddc46824 [Auto Sync] Update server_args.py (20251223) (#15700)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
2025-12-24 14:55:14 -08:00