fzyzcjy
|
a368df2818
|
Tiny add more error info for bench_serving (#14827)
|
2025-12-11 17:25:49 +08:00 |
|
Alison Shao
|
f85460fb19
|
Avoid deleting entire cache for missing shards (#14754 follow-up) (#14853)
|
2025-12-11 01:17:04 -08:00 |
|
Douglas Yang
|
e52cf30e81
|
fix: adding temporary bypass for nightly tests (#14876)
|
2025-12-11 00:09:03 -08:00 |
|
Shangming Cai
|
a076d75e98
|
[Fix] Remove unused import from test_disaggregation_hicache.py (#14880)
|
2025-12-11 15:50:47 +08:00 |
|
shicanwei.scw
|
8348725d9e
|
[CI][BUG] fix ib setup for disaggregation hicache test (#14877)
Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
|
2025-12-11 15:40:48 +08:00 |
|
Kangyan-Zhou
|
2856624156
|
Only count limitations for previous runs that reaches the test stages (#14856)
|
2025-12-10 23:17:34 -08:00 |
|
Yuhao Yang
|
b62fe8504c
|
fix nightly vlm ci : restore original eval for requests without regex (#14875)
|
2025-12-10 23:13:25 -08:00 |
|
Tiance Wang
|
624725cb5e
|
Move and update MindSpore docs, make it appear on the online documentation (#14861)
Co-authored-by: wangtiance <tiancew@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-10 23:03:50 -08:00 |
|
Douglas Yang
|
1a96e66493
|
fix: creating blobs only once for publish trace retries (#14845)
|
2025-12-10 22:40:06 -08:00 |
|
Kangyan-Zhou
|
32829b1638
|
Remove myself to test CI gate issue (#14871)
|
2025-12-10 22:37:01 -08:00 |
|
yuchengz816-bot
|
e54307f26a
|
[6/n] Fix num_token_non_padded computation in prefill (#14313)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Runkai Tao <rt572@physics.rutger.edu>
|
2025-12-10 19:15:19 -08:00 |
|
Trang Do
|
8642dbe416
|
Refactor Marlin MoeRunner (#14554)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-12-10 19:15:00 -08:00 |
|
Baizhou Zhang
|
7dcad45cad
|
[CI] Temp disable gb200 test (#14865)
|
2025-12-10 19:05:16 -08:00 |
|
roikoren755
|
7c985331dd
|
Enable TP for Mamba-based models (#14811)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2025-12-10 18:03:50 -08:00 |
|
Liangsheng Yin
|
bd7824b24d
|
Minor code style fix for dllm (#14836)
|
2025-12-11 10:35:42 +09:00 |
|
Binyao Jiang
|
312df1d6c0
|
Fix TestGLM41VPPAccuracy test flakiness (#14848)
|
2025-12-10 16:59:58 -08:00 |
|
Lianmin Zheng
|
25e97380e3
|
Fix CUDA version handling in ci_install_deepep.sh (#14854)
|
2025-12-10 16:52:30 -08:00 |
|
Alison Shao
|
b6523a4f72
|
fix: restrict cache validation behaviors to CI only (#14849)
|
2025-12-10 16:03:53 -08:00 |
|
b8zhong
|
c51efb8b84
|
fix fp8 gemm nightly CI (#14844)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
|
2025-12-10 15:57:51 -08:00 |
|
Simo Lin
|
bcc5483ed7
|
[model-gateway] fix import order in oai conversation (#14851)
|
2025-12-10 15:36:42 -08:00 |
|
Simo Lin
|
ccf2602773
|
[model-gateway] code clean up on oai router (#14850)
|
2025-12-10 15:34:27 -08:00 |
|
Binyao Jiang
|
a4992873d4
|
Treat unittest SkipTest exception as pass instead of as failure (#14847)
|
2025-12-10 15:28:21 -08:00 |
|
michael-amd
|
c97ce39181
|
[AMD] Add model to AMD nightly test (#14442)
|
2025-12-10 11:59:44 -08:00 |
|
Simo Lin
|
c032b559a8
|
[model-gateway] adds default implementations to RouterTrait in mod.rs (#14841)
|
2025-12-10 11:46:35 -08:00 |
|
b8zhong
|
da9b801eb7
|
fix lora target all + csgmv backend (#14796)
|
2025-12-10 11:42:51 -08:00 |
|
Siyuan Chen
|
0e54a69548
|
[bugfix] qwen25-VL support lora (#14638)
|
2025-12-10 11:38:51 -08:00 |
|
Praneth Paruchuri
|
e99ee0c695
|
[model-gateway] Fix incompatible metric comparison in PowerOfTwo policy (#14823)
|
2025-12-10 11:15:38 -08:00 |
|
Yuhao Yang
|
c1bd5ee8c5
|
Revert transformers to 4.57.1 (#14801)
|
2025-12-10 11:04:36 -08:00 |
|
Yineng Zhang
|
ef1ab2302a
|
[Auto Sync] Update tool_chat_template_deepseekv31.jinja (20251210) (#14837)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Jue Wang <zjuwangjue@gmail.com>
|
2025-12-10 10:56:24 -08:00 |
|
Mick
|
d659873762
|
[diffusion] CI: use unified sampling_params for CI (#14045)
|
2025-12-11 01:18:56 +08:00 |
|
Mick
|
6c5ebc0ef7
|
[diffusion] parallel: pad tokens for video models under sp (#14833)
|
2025-12-11 01:15:37 +08:00 |
|
Ke Bao
|
5b5571a8da
|
Apply back moe_sum_reduce for fused_marlin_moe (#14829)
|
2025-12-11 00:39:41 +08:00 |
|
Alison Shao
|
1698c2341b
|
[CI] Reduce stage-b auto-partition from 4 to 2 (#14769)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-11 01:32:33 +09:00 |
|
fzyzcjy
|
3d82c0f17f
|
[model-gateway] support engine response http status statistics in router (#14712)
|
2025-12-10 08:30:53 -08:00 |
|
fzyzcjy
|
617e9b3bf8
|
[model-gateway] support customizing Prometheus duration buckets (#14716)
|
2025-12-10 08:30:03 -08:00 |
|
Li Jinliang
|
83e35a7c29
|
[diffusion] doc: fix tiny typo in multimodal_gen/README.md (#14830)
|
2025-12-11 00:27:25 +08:00 |
|
Simo Lin
|
2543666c8e
|
[model-gateway] add anthropic message api spec (#14834)
|
2025-12-10 08:26:51 -08:00 |
|
Alison Shao
|
f732f8ea57
|
Fix CI registry scan to only check test/registered directory (#14812)
|
2025-12-11 01:09:03 +09:00 |
|
Liangsheng Yin
|
503880dbbe
|
[CI] fix UT success check in test_eagle_infer_beta_dp_attention.py (#14831)
|
2025-12-11 01:00:50 +09:00 |
|
liupeng374
|
b8cfa02c01
|
[NPU] bug fix for mtp and w4a8 (#14806)
|
2025-12-10 22:55:13 +08:00 |
|
fzyzcjy
|
d85fecb536
|
Tiny fix incorrect worker removal command (#14822)
|
2025-12-10 22:29:39 +08:00 |
|
fzyzcjy
|
6634f67bd5
|
Fix router keep nonzero metrics after worker is deleted (#14819)
|
2025-12-10 22:29:30 +08:00 |
|
fzyzcjy
|
5eccaf7737
|
Support HTTP response status code prometheus metrics (#14710)
|
2025-12-10 22:28:57 +08:00 |
|
Wenyi Xu
|
d7f6320bb8
|
[model-gateway] Dynamically Populate Tool Call Parser Choices (#14807)
|
2025-12-10 06:04:22 -08:00 |
|
ybyang
|
766476f52a
|
[SMG-GO] implement a Go SGLang Model Gateway - OpenAI Compatible API Server (#14770)
|
2025-12-10 06:03:28 -08:00 |
|
Xiaoyu Zhang
|
12b7a4fab0
|
[diffusion] performance: refactor diffusion fuse qkv and apply to qwen-image (#14793)
|
2025-12-10 18:55:41 +08:00 |
|
Yuhao Yang
|
02f1e81e2d
|
Revert "fix: checking if tokenizer is in cache before downloading from HF" (#14808)
|
2025-12-10 01:14:35 -08:00 |
|
Prozac614
|
908c7186af
|
[diffusion] CI: Add LoRA support to diffusion server configuration and test cases (#14697)
|
2025-12-10 16:51:47 +08:00 |
|
Yuan Luo
|
03836d85d2
|
[GLM-4.6V] Support Pipeline Parallelism for GLM-4.6V & GLM-4.1V (#14720)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-12-10 16:40:12 +08:00 |
|
Mick
|
87dbdddc93
|
[diffusion] profile: early exit when enough steps are captured to reduce the size of the trace file (#14803)
|
2025-12-10 16:11:22 +08:00 |
|