Commit Graph

7701 Commits

Author SHA1 Message Date
fzyzcjy
a368df2818 Tiny add more error info for bench_serving (#14827) 2025-12-11 17:25:49 +08:00
Alison Shao
f85460fb19 Avoid deleting entire cache for missing shards (#14754 follow-up) (#14853) 2025-12-11 01:17:04 -08:00
Douglas Yang
e52cf30e81 fix: adding temporary bypass for nightly tests (#14876) 2025-12-11 00:09:03 -08:00
Shangming Cai
a076d75e98 [Fix] Remove unused import from test_disaggregation_hicache.py (#14880) 2025-12-11 15:50:47 +08:00
shicanwei.scw
8348725d9e [CI][BUG] fix ib setup for disaggregation hicache test (#14877)
Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
2025-12-11 15:40:48 +08:00
Kangyan-Zhou
2856624156 Only count limitations for previous runs that reaches the test stages (#14856) 2025-12-10 23:17:34 -08:00
Yuhao Yang
b62fe8504c fix nightly vlm ci : restore original eval for requests without regex (#14875) 2025-12-10 23:13:25 -08:00
Tiance Wang
624725cb5e Move and update MindSpore docs, make it appear on the online documentation (#14861)
Co-authored-by: wangtiance <tiancew@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-10 23:03:50 -08:00
Douglas Yang
1a96e66493 fix: creating blobs only once for publish trace retries (#14845) 2025-12-10 22:40:06 -08:00
Kangyan-Zhou
32829b1638 Remove myself to test CI gate issue (#14871) 2025-12-10 22:37:01 -08:00
yuchengz816-bot
e54307f26a [6/n] Fix num_token_non_padded computation in prefill (#14313)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Runkai Tao <rt572@physics.rutger.edu>
2025-12-10 19:15:19 -08:00
Trang Do
8642dbe416 Refactor Marlin MoeRunner (#14554)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
2025-12-10 19:15:00 -08:00
Baizhou Zhang
7dcad45cad [CI] Temp disable gb200 test (#14865) 2025-12-10 19:05:16 -08:00
roikoren755
7c985331dd Enable TP for Mamba-based models (#14811)
Signed-off-by: Roi Koren <roik@nvidia.com>
2025-12-10 18:03:50 -08:00
Liangsheng Yin
bd7824b24d Minor code style fix for dllm (#14836) 2025-12-11 10:35:42 +09:00
Binyao Jiang
312df1d6c0 Fix TestGLM41VPPAccuracy test flakiness (#14848) 2025-12-10 16:59:58 -08:00
Lianmin Zheng
25e97380e3 Fix CUDA version handling in ci_install_deepep.sh (#14854) 2025-12-10 16:52:30 -08:00
Alison Shao
b6523a4f72 fix: restrict cache validation behaviors to CI only (#14849) 2025-12-10 16:03:53 -08:00
b8zhong
c51efb8b84 fix fp8 gemm nightly CI (#14844)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
2025-12-10 15:57:51 -08:00
Simo Lin
bcc5483ed7 [model-gateway] fix import order in oai conversation (#14851) 2025-12-10 15:36:42 -08:00
Simo Lin
ccf2602773 [model-gateway] code clean up on oai router (#14850) 2025-12-10 15:34:27 -08:00
Binyao Jiang
a4992873d4 Treat unittest SkipTest exception as pass instead of as failure (#14847) 2025-12-10 15:28:21 -08:00
michael-amd
c97ce39181 [AMD] Add model to AMD nightly test (#14442) 2025-12-10 11:59:44 -08:00
Simo Lin
c032b559a8 [model-gateway] adds default implementations to RouterTrait in mod.rs (#14841) 2025-12-10 11:46:35 -08:00
b8zhong
da9b801eb7 fix lora target all + csgmv backend (#14796) 2025-12-10 11:42:51 -08:00
Siyuan Chen
0e54a69548 [bugfix] qwen25-VL support lora (#14638) 2025-12-10 11:38:51 -08:00
Praneth Paruchuri
e99ee0c695 [model-gateway] Fix incompatible metric comparison in PowerOfTwo policy (#14823) 2025-12-10 11:15:38 -08:00
Yuhao Yang
c1bd5ee8c5 Revert transformers to 4.57.1 (#14801) 2025-12-10 11:04:36 -08:00
Yineng Zhang
ef1ab2302a [Auto Sync] Update tool_chat_template_deepseekv31.jinja (20251210) (#14837)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Jue Wang <zjuwangjue@gmail.com>
2025-12-10 10:56:24 -08:00
Mick
d659873762 [diffusion] CI: use unified sampling_params for CI (#14045) 2025-12-11 01:18:56 +08:00
Mick
6c5ebc0ef7 [diffusion] parallel: pad tokens for video models under sp (#14833) 2025-12-11 01:15:37 +08:00
Ke Bao
5b5571a8da Apply back moe_sum_reduce for fused_marlin_moe (#14829) 2025-12-11 00:39:41 +08:00
Alison Shao
1698c2341b [CI] Reduce stage-b auto-partition from 4 to 2 (#14769)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2025-12-11 01:32:33 +09:00
fzyzcjy
3d82c0f17f [model-gateway] support engine response http status statistics in router (#14712) 2025-12-10 08:30:53 -08:00
fzyzcjy
617e9b3bf8 [model-gateway] support customizing Prometheus duration buckets (#14716) 2025-12-10 08:30:03 -08:00
Li Jinliang
83e35a7c29 [diffusion] doc: fix tiny typo in multimodal_gen/README.md (#14830) 2025-12-11 00:27:25 +08:00
Simo Lin
2543666c8e [model-gateway] add anthropic message api spec (#14834) 2025-12-10 08:26:51 -08:00
Alison Shao
f732f8ea57 Fix CI registry scan to only check test/registered directory (#14812) 2025-12-11 01:09:03 +09:00
Liangsheng Yin
503880dbbe [CI] fix UT success check in test_eagle_infer_beta_dp_attention.py (#14831) 2025-12-11 01:00:50 +09:00
liupeng374
b8cfa02c01 [NPU] bug fix for mtp and w4a8 (#14806) 2025-12-10 22:55:13 +08:00
fzyzcjy
d85fecb536 Tiny fix incorrect worker removal command (#14822) 2025-12-10 22:29:39 +08:00
fzyzcjy
6634f67bd5 Fix router keep nonzero metrics after worker is deleted (#14819) 2025-12-10 22:29:30 +08:00
fzyzcjy
5eccaf7737 Support HTTP response status code prometheus metrics (#14710) 2025-12-10 22:28:57 +08:00
Wenyi Xu
d7f6320bb8 [model-gateway] Dynamically Populate Tool Call Parser Choices (#14807) 2025-12-10 06:04:22 -08:00
ybyang
766476f52a [SMG-GO] implement a Go SGLang Model Gateway - OpenAI Compatible API Server (#14770) 2025-12-10 06:03:28 -08:00
Xiaoyu Zhang
12b7a4fab0 [diffusion] performance: refactor diffusion fuse qkv and apply to qwen-image (#14793) 2025-12-10 18:55:41 +08:00
Yuhao Yang
02f1e81e2d Revert "fix: checking if tokenizer is in cache before downloading from HF" (#14808) 2025-12-10 01:14:35 -08:00
Prozac614
908c7186af [diffusion] CI: Add LoRA support to diffusion server configuration and test cases (#14697) 2025-12-10 16:51:47 +08:00
Yuan Luo
03836d85d2 [GLM-4.6V] Support Pipeline Parallelism for GLM-4.6V & GLM-4.1V (#14720)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-12-10 16:40:12 +08:00
Mick
87dbdddc93 [diffusion] profile: early exit when enough steps are captured to reduce the size of the trace file (#14803) 2025-12-10 16:11:22 +08:00