gaopengff
|
7541da15d2
|
Fix prefill latency performance drop of bench serving (#14592)
|
2026-01-29 21:28:17 -08:00 |
|
Douglas Yang
|
4c7136bb36
|
feature: adding openai compatible API request to bench_serving (#17219)
|
2026-01-23 16:04:28 -08:00 |
|
fzyzcjy
|
d874c8bba4
|
Tiny support http headers in bench serving (#16606)
|
2026-01-07 10:15:17 +08:00 |
|
fzyzcjy
|
ed307a40bf
|
Tiny add --gsp-ordered (#16575)
|
2026-01-06 22:43:52 +08:00 |
|
Raayan Dhar
|
84d13c54bb
|
feat: only add input vision tokens in bench_serving result if vision dataset is used (#15492)
Signed-off-by: Raayan Dhar raayan.dhar@gmail.com <raayan.dhar@gmail.com>
|
2026-01-06 12:14:42 +08:00 |
|
fzyzcjy
|
c105a3124b
|
Support multi-round conversations in bench_serving (#6135)
|
2026-01-06 11:59:39 +08:00 |
|
Kangyan-Zhou
|
12b89e51d8
|
Add P90/99 e2e latency in bench_serving script (#16245)
|
2025-12-31 15:40:33 -08:00 |
|
jiapingW
|
278012caa0
|
[Feature] support bench jsonl files with sharegpt format (#15057)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-12-29 10:06:11 -08:00 |
|
fzyzcjy
|
c12567274a
|
Tiny change bench-serving to use routing key header (#15827)
|
2025-12-25 21:29:49 +08:00 |
|
fzyzcjy
|
ad7f35fb9d
|
Support gsp send routing id in bench serving (#15588)
|
2025-12-22 11:32:24 +08:00 |
|
fzyzcjy
|
f4100732b8
|
Tiny fix bench serving GSP mode cache file strategy (#15587)
|
2025-12-22 11:27:55 +08:00 |
|
Junrong Lin
|
bed301a5ac
|
[Feature] Enable return routed experts (#12162)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-21 15:16:43 +08:00 |
|
b8zhong
|
e72b02db28
|
tiny fix lint on main (#15424)
|
2025-12-18 15:55:02 -08:00 |
|
Frank
|
9749d3e346
|
Update benchmarks to use HF token from environment. (#15421)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-18 13:47:27 -08:00 |
|
Tianyu Guo
|
9acb21ae27
|
feat: support EPD disaggregation (#12263)
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
Co-authored-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: Nicholas <45984215+liusy58@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
|
2025-12-14 22:30:08 +08:00 |
|
fzyzcjy
|
df111bc0fe
|
Super tiny add gsp-fast-prepare (#14992)
|
2025-12-13 09:45:21 +08:00 |
|
fzyzcjy
|
45eeeb9a3c
|
Tiny support range ratio in GSP in bench serving (#14828)
|
2025-12-11 17:26:07 +08:00 |
|
fzyzcjy
|
a368df2818
|
Tiny add more error info for bench_serving (#14827)
|
2025-12-11 17:25:49 +08:00 |
|
fzyzcjy
|
119fd956fb
|
Tiny support printing requests in bench_serving for observability (#14652)
|
2025-12-09 08:27:58 +08:00 |
|
Baizhou Zhang
|
42fcf5438f
|
Revert "tiny remove deprecated endpoint call" (#14533)
|
2025-12-05 23:48:54 -08:00 |
|
b8zhong
|
ec7b2c16d9
|
tiny remove deprecated endpoint call (#13607)
|
2025-12-05 09:54:49 -08:00 |
|
Xiaoyu Zhang
|
9c80072845
|
Add peak output tokens per second in bench_serving (#14165)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-12-01 17:47:54 +08:00 |
|
fzyzcjy
|
487c8d4df3
|
Tiny add several args to bench serving (#14181)
|
2025-12-01 14:45:47 +08:00 |
|
Binyao Jiang
|
90c18a16cb
|
[GLM4.6v] Required changes for bumping up to transformer 5.x (#13229)
|
2025-11-18 10:58:00 +08:00 |
|
Douglas Yang
|
7ee3e36412
|
Fix: test_vlm_offline_throughput output throughput (#13279)
|
2025-11-15 14:31:48 +08:00 |
|
Ling Zhang
|
0050ff254f
|
[BugFix] fix bench_serving error when multimodal image is testing (#13254)
|
2025-11-14 22:46:19 +08:00 |
|
Rohan Potdar
|
c9b581644d
|
Dump total_throughput to output-file in bench_serving.py (#9790)
|
2025-11-12 18:48:28 +08:00 |
|
Ziming Huang
|
e8114102fa
|
[Fix] Update text_chunks in bench_serving chat completions (#13041)
|
2025-11-11 11:41:11 -08:00 |
|
elvischenv
|
71aea45c41
|
[Fix] Add TPOT back to bench_serving (#12976)
|
2025-11-11 17:32:24 +08:00 |
|
yinghui
|
b88fab3111
|
fix: add seed bench_serving to cache key, remove redundant function definition. (#12680)
|
2025-11-04 23:39:11 -08:00 |
|
Glen Liu
|
cbf23dbbfa
|
[Feature] add --lora-request-distribution arg to bench_serving.py and support skewed and distinct workloads (#12175)
|
2025-11-04 21:41:40 -08:00 |
|
fzyzcjy
|
8834260739
|
Super tiny dump server info such as args in bench for post analysis (#12550)
|
2025-11-03 14:24:08 -08:00 |
|
fzyzcjy
|
fd7a72d62d
|
Super tiny allow profile activities in bench_serving (#12549)
|
2025-11-03 14:23:18 -08:00 |
|
Yi Zhang
|
21a8fa16ea
|
tiny optimize for bench serving (#12553)
|
2025-11-03 14:13:18 -08:00 |
|
Ke Bao
|
f7f9e41b36
|
Fix run benchmark (#12473)
|
2025-11-01 02:39:48 +08:00 |
|
daniel, chen
|
cf0c24150a
|
add served model name in bench serving (#12428)
|
2025-11-01 01:28:11 +08:00 |
|
Zaili Wang
|
50b6842b4b
|
fix: Add default value for backend in sample_mmmu_requests (#12256)
|
2025-10-31 19:31:40 +08:00 |
|
fzyzcjy
|
fb52d35f63
|
Super tiny fix AMD ci (#12378)
|
2025-10-29 23:25:18 -07:00 |
|
fzyzcjy
|
25c5049870
|
Super tiny add tag for benchmark scripts (#12340)
|
2025-10-30 11:19:14 +08:00 |
|
Liangsheng Yin
|
ea385ae85a
|
Fix ITL metrics when using openai endpoint with spec (#12156)
|
2025-10-26 18:06:25 +08:00 |
|
Xinyuan Tong
|
b9fb74f3bc
|
fix: bench_serving ITL calculation when using spec-decoding (#12064)
|
2025-10-24 17:02:44 +08:00 |
|
Zaili Wang
|
92009bd28e
|
fix: fix MMMU loading issue (#11759)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-10-23 20:21:38 -07:00 |
|
b8zhong
|
d0a64c7e2c
|
vlm: enforce pybase64 for image and str encode/decode (#10700)
|
2025-10-21 19:05:32 +08:00 |
|
Xiaoyu Zhang
|
88a6f9dab5
|
bench_serving support PD Disaggregation (#11542)
|
2025-10-13 19:43:26 -07:00 |
|
Mick
|
0c0779d667
|
ci: improve nightly-ci (#11385)
|
2025-10-12 21:19:34 -07:00 |
|
shaharmor98
|
fa7e2c3049
|
fix bench_serving mishandling of internal states (#11376)
Signed-off-by: Shahar Mor <smor@nvidia.com>
|
2025-10-09 19:24:50 +08:00 |
|
Kevin Xiang Li
|
e3bb7f5ae6
|
benchmark: enhance configurable multimodal benchmarking in bench_serving (#9812)
Co-authored-by: Xiang (Kevin) Li <lik@nvidia.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-10-08 01:31:36 -07:00 |
|
Lifu Huang
|
92473e2e34
|
Support LoRA in bench_serving oai interface (#11318)
|
2025-10-08 01:28:58 -07:00 |
|
fzyzcjy
|
fdc4e1e570
|
Tiny move files to utils folder (#11166)
|
2025-10-03 22:40:06 +08:00 |
|
Muqi Li
|
51f7c6bd3c
|
Add auth to get server info (#10751)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-09-27 02:54:39 -07:00 |
|