Commit Graph
13 Commits
Author SHA1 Message Date
Baizhou Zhang 42fcf5438f Revert "tiny remove deprecated endpoint call" (#14533) 2025-12-05 23:48:54 -08:00
b8zhong ec7b2c16d9 tiny remove deprecated endpoint call (#13607) 2025-12-05 09:54:49 -08:00
Lzhang-hub 2847e5c4b4 fix bench_speculative bug (#13197) 2025-11-20 17:09:04 +08:00
Xiaoyu Zhang 8b5e2c5368 [Tiny fix] Fix bench_speculative.py run bug (#13416) 2025-11-17 18:58:19 +08:00
Zaili Wang 50b6842b4b fix: Add default value for backend in sample_mmmu_requests (#12256) 2025-10-31 19:31:40 +08:00
Lzhang-hub 4efe2c57c9 support vlm model spec bench (#10173) 2025-09-10 13:37:04 +08:00
Kay Yan 975a5ec69c [fix] update bench_speculative.py for compatibility (#7764)
Signed-off-by: Kay Yan <kay.yan@daocloud.io>
2025-07-04 16:32:54 +08:00
Yineng Zhang 7282ab741a fix: update bench_speculative (#5649) 2025-04-22 16:08:15 -07:00
Baizhou Zhang 6fb29ffd9e Deprecate enable-flashinfer-mla and enable-flashmla (#5480) 2025-04-17 01:43:33 -07:00
lukecandyinfan98 a53fe428f9 Support FlashMLA backend (#4472)
Co-authored-by: yinfan98 <1106310035@qq.com>
2025-03-16 09:07:06 -07:00
Ke Bao f1d09a6541 Update bench speculative script (#4235) 2025-03-09 12:19:01 -07:00
Lianmin Zheng 935cda944b Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00
ac2387279e Support penalty in overlap mode; return logprob with chunked prefill; improve benchmark scripts (#3988)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: dhou-xai <dhou@x.ai>
Co-authored-by: Hanming Lu <hanming_lu@berkeley.edu>
2025-03-03 00:12:04 -08:00