Liangsheng Yin
|
4d3976b6c5
|
[HiCache] Check in-flight async ops in is_fully_idle() before attach/detach (#20746)
|
2026-03-17 17:28:26 -07:00 |
|
Bruce Wu
|
70a6fb53af
|
Enable embedding lookup/lora_a logic for chunked backend (#17692)
Co-authored-by: Bruce Wu <mogicianwu@fb.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com>
|
2026-03-16 11:37:58 -07:00 |
|
Sugar920
|
895e56097c
|
Add NPU basic function testcases (#19382)
Co-authored-by: cy <chenyang08056032@163.com>
Co-authored-by: Cherry_ming <136634645@qq.com>
|
2026-03-16 15:09:56 +08:00 |
|
Liangsheng Yin
|
f0458e0b49
|
[Utils] Move network/socket utilities from common.py to network.py (#20646)
|
2026-03-15 20:35:24 -07:00 |
|
SoluMilken
|
c95dc88f86
|
[CI] migrate ascend-gptq from test/srt to test/registered (#19628)
|
2026-03-14 00:28:57 -07:00 |
|
Pai Liu
|
65dd08153d
|
Fix Test* mixin classes being collected as standalone pytest tests (#20417)
|
2026-03-12 18:18:45 -07:00 |
|
roikoren755
|
067353f67b
|
[Test] Refactor KL divergence and prefix cache branching to kits (#19715)
|
2026-03-12 16:11:59 +08:00 |
|
Артем Савкин
|
ed42af99a9
|
[NPU] [Quantization] w4a4 MoE layer support (#18924)
|
2026-03-11 16:52:35 +03:00 |
|
Jacob0226
|
dadd4dde83
|
[AMD] Skip the flaky test for lora ci test. (#20175)
Co-authored-by: YC Tseng <yctseng@amd.com>
|
2026-03-09 23:15:30 -07:00 |
|
shuwenn
|
5a11ae19c1
|
[CI] fix: notebook ci often OOM (#20199)
|
2026-03-09 22:32:41 -07:00 |
|
Mohammad Miadh Angkad
|
ca997b7ba9
|
Add min_p and chat-template kwargs support to run_eval (#19571)
|
2026-03-09 14:53:09 -07:00 |
|
Ke Bao
|
2e444bdced
|
Move stop words to args in send one (#20193)
|
2026-03-09 23:05:32 +08:00 |
|
Xinyuan Tong
|
4a757990a1
|
[VLM] Replace decord with torchcodec for video decoding (#20055)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: BakerBunker <17872844+BakerBunker@users.noreply.github.com>
|
2026-03-09 19:23:49 +08:00 |
|
Alison Shao
|
0f62da6953
|
[CI] Show test partition assignments after checkout (#20085)
Co-authored-by: Alison Shao <alisonshao@mac.lan>
|
2026-03-07 13:50:49 -08:00 |
|
Ajay Anubolu
|
13af7cbb02
|
fix: use consistent time denominator for throughput metrics in bench_one_batch_server (#19223)
|
2026-03-05 15:58:17 -08:00 |
|
kpham-sgl
|
346a4131cf
|
[Spec] Refactor NaN/OOB checks to async maybe_detect_* with env-var control (#19899)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-03-05 13:51:05 -08:00 |
|
Mohammad Miadh Angkad
|
41fd53fe37
|
Fix profile_activities parameter name in bench_one_batch_server_internal.py (#19954)
|
2026-03-05 10:34:06 -08:00 |
|
Kalyan Kumar
|
c1df359b44
|
Add XPU profiler activity support in benchmark code (#12981)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-04 23:22:56 -08:00 |
|
Jonah Bernard
|
fb37c0a400
|
[args] Add Expert Parallelism Argument To SRT Runner (#18492)
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
|
2026-03-03 14:16:35 -08:00 |
|
Kangyan-Zhou
|
dc92f88a21
|
Enhance bench_multiturn.py with OpenAI API support and richer metrics (#19724)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-03-03 13:48:04 -08:00 |
|
almaslof
|
b0f26698f5
|
feat(benchmark script): add similar to vllm --ready-check-timeout-sec parameter (#15466)
|
2026-03-03 13:44:38 -08:00 |
|
Glen Liu
|
cc860a2198
|
[TestFix] change LoRA tests to use NVIDIA adapter instead of Nutanix (#19642)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-02 12:55:41 -08:00 |
|
Henry
|
e5edf222cd
|
[WIP]enable mxfp8 on nvidia sm120 (#19112)
Co-authored-by: Your Name <you@example.com>
|
2026-03-01 19:06:43 -08:00 |
|
yrk111222
|
e6da514c2c
|
CI: use 'sglang serve' in CI tests (#18597)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
|
2026-02-27 14:00:41 -08:00 |
|
Alison Shao
|
c2dce06d9f
|
Fix parallel tool call test for speculative decoding variants (#19370)
Co-authored-by: Alison Shao <alisonshao@Mac.attlocal.net>
Co-authored-by: Alison Shao <alisonshao@MacBook-Pro-D2W773R9CD.local>
|
2026-02-26 18:31:20 -08:00 |
|
billishyahao
|
60eeef7370
|
[AMD][with CI Fix] support two batch overlapping for mori ep (#19216)
Co-authored-by: Duyi-Wang <duyi.wang@amd.com>
Co-authored-by: kkHuang-amd <wunhuang@amd.com>
Co-authored-by: Feiyue Zhai <feiyue.zhai@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-02-25 02:14:08 -08:00 |
|
Julian Huang
|
a55f658835
|
[Misc] Normalize --host parameter to use plain hostname without scheme (#19309)
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-02-25 00:37:24 -08:00 |
|
Ratish P
|
ae6f6e1495
|
[Refactor] Benchmark: Add typed DatasetArgs/Loader registry and CPU dataset unit tests (#19147)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2026-02-24 12:22:01 -08:00 |
|
Liangsheng Yin
|
9fac90d85d
|
[CI] Tiny enhance the dp attention load blance benchmark (#19194)
|
2026-02-23 14:33:32 -08:00 |
|
Liangsheng Yin
|
2aa3fe394e
|
[CI] fix the teardown output of disaggregation test (#19193)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-02-23 12:41:03 -08:00 |
|
Qiaolin Yu
|
42b1019881
|
Fix bench_one_batch_server by moving the print statements (#19175)
|
2026-02-22 22:06:25 -08:00 |
|
Baizhou Zhang
|
fa80b9beba
|
[CI] Skip some subtests for tool call parser (#19172)
|
2026-02-23 12:20:12 +08:00 |
|
Baizhou Zhang
|
43f83525c0
|
Revert "[AMD] support two batch overlapping for mori ep #17953" (#19161)
|
2026-02-23 01:19:23 +08:00 |
|
Liangsheng Yin
|
1f2da824dd
|
[Benchmark] Remove re-exports from bench_serving.py (#19130)
|
2026-02-21 14:30:30 -08:00 |
|
Lianmin Zheng
|
2928dfb8fa
|
[Auto Sync] Update bench_one_batch_server_internal.py (20260221) (#19097)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
|
2026-02-20 18:19:14 -08:00 |
|
Qiaolin Yu
|
96bae2355e
|
Add generated-shared-prefix dataset in bench_one_batch (#18986)
|
2026-02-20 13:33:10 -08:00 |
|
billishyahao
|
fbb6098487
|
[AMD] support two batch overlapping for mori ep (#17953)
Co-authored-by: kkHuang-amd <wunhuang@amd.com>
Co-authored-by: Feiyue Zhai <feiyue.zhai@amd.com>
Co-authored-by: Duyi-Wang <duyi.wang@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-02-20 08:45:55 -08:00 |
|
Alison Shao
|
34d975b18f
|
Fix eval tests not capturing server launch failures (#18886)
|
2026-02-18 07:59:03 +08:00 |
|
Lianmin Zheng
|
e02a9bec8d
|
Refactor sampler: Use a better hash function for deterministic sampling and clear dispatch for probs/logprobs/logits sampling paths (#18915)
Co-authored-by: Sehoon Kim <sehoon@x.ai>
|
2026-02-17 15:41:23 -08:00 |
|
Liangsheng Yin
|
83a475e8d7
|
feat: add cuda core dump CI warpper (#18909)
|
2026-02-17 14:49:26 -08:00 |
|
Alison Shao
|
86c181e335
|
Fix test_lora_qwen3 nightly failure: replace adapter with added_tokens (#18884)
|
2026-02-16 14:35:06 +08:00 |
|
fzyzcjy
|
4c7f986c6b
|
Extract dumper and prefill delayer tests common utils (#18857)
|
2026-02-15 18:33:23 +08:00 |
|
Lianmin Zheng
|
8b2020584c
|
[Auto Sync] Update test_deterministic.py (20260214) (#18839)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Jiayi Yuan <34369239+jy-yuan@users.noreply.github.com>
|
2026-02-14 17:19:30 -08:00 |
|
Kangyan-Zhou
|
ae95869292
|
Enable SGLANG_ENABLE_SPEC_V2 for nightly speculative decoding tests (#18719)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-02-14 23:00:33 +08:00 |
|
Ke Bao
|
f51e9d9ca1
|
Add ci test for ring model (#18829)
|
2026-02-14 22:20:23 +08:00 |
|
Liangsheng Yin
|
dcea74d63f
|
Add timeout abort kits for normal / eagle. (#18815)
|
2026-02-13 17:57:30 -08:00 |
|
shuwenn
|
3299c4f9c1
|
[CI] feat: add early exit to wait_for_server when process dies (#18602)
|
2026-02-13 16:46:09 -08:00 |
|
Hudson Xing
|
f3656432c7
|
add tool_choice=auto nightly test case (#18302)
|
2026-02-12 19:28:05 +08:00 |
|
YC Tseng
|
20554a0a4f
|
[AMD] rocm 7.2 image release, PR test, Nightly Test (#17799)
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
|
2026-02-11 21:29:25 -08:00 |
|
Liangsheng Yin
|
cd90346a2b
|
Add cache hit rate UT (#18566)
|
2026-02-10 21:27:41 -08:00 |
|