Commit Graph
611 Commits
Author SHA1 Message Date
Lianmin Zheng 048685430d Improve process creation (#1534) 2024-09-29 02:36:12 -07:00
Liangsheng Yin fd9ad817ec Organize image inputs (#1531) 2024-09-29 06:28:55 +00:00
Lianmin Zheng e165a9fc1b Make detokenizer_manager.py not asyncio (#1532) 2024-09-28 19:33:09 -07:00
Ying Sheng 9aa6553d2a [Feature] Support reward model LxzGordon/URM-LLaMa-3.1-8B (#1525) 2024-09-27 23:32:11 -07:00
Lianmin Zheng 2854a5ea9f Fix the overhead due to penalizer in bench_latency (#1496) 2024-09-23 07:38:14 -07:00
2a99993cd9 Pr fix max workers (#1456)
Co-authored-by: baolujia <baolujia@shizhuang-inc.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2024-09-22 02:20:26 -07:00
Xiao Yu 5752f25eef Fixed n>1 causing list index out of range with VLM (#1449) 2024-09-18 00:46:32 -07:00
Liangsheng Yin 7c162fa9c5 Fix schedule bug (#1451) 2024-09-17 22:59:32 -07:00
Liangsheng Yin 36078fb247 fix schedule bug (#1450) 2024-09-17 16:33:53 -07:00
Ke Bao c6b6d2e71b Enable MLA by default (#1447) 2024-09-17 11:42:48 +00:00
Lianmin Zheng 2fa5cec775 Simplify sampler and its error handling (#1441) 2024-09-16 21:23:31 -07:00
Lianmin Zheng 27b557aea7 Clean up model loader (#1440) 2024-09-16 18:16:27 -07:00
zifeitong 93dffd699b Add constrained_json_whitespace_pattern to ServerArgs (#1438) 2024-09-16 13:29:18 -07:00
Lianmin Zheng 9ba1f09760 [Fix] Fix logprob and normalized_logprob (#1428) 2024-09-15 06:36:06 -07:00
Liangsheng Yin 70b6802982 Optimize conflicts between CUDA graph and vocab mask tensors (#1392) 2024-09-13 20:27:53 -07:00
Lianmin Zheng b912de11b0 Make stop reason a dict instead of str (#1407) 2024-09-12 20:47:31 -07:00
Ying Sheng 712216928f [Feature] Initial support for multi-LoRA serving (#1307) 2024-09-12 16:46:14 -07:00
Lianmin Zheng fec185ce0c Refactor attention backend (#1381) 2024-09-11 11:44:26 -07:00
Liangsheng Yin 144bc70fcc Organize flashinfer indices update (#1378) 2024-09-10 17:38:59 -07:00
Lianmin Zheng 46094e0c1b Deprecate --disable-flashinfer and introduce --attention-backend (#1380) 2024-09-10 17:11:16 -07:00
Lianmin Zheng 3a6e8b6d78 [Minor] move triton attention kernels into a separate folder (#1379) 2024-09-10 15:15:08 -07:00
Liangsheng Yin fbb4754cb8 Fix vocab mask update bug (#1376) 2024-09-10 13:10:36 -07:00
wangchao fec2d1223c [Fix] fix bug of undefined is_single in meth create_abort_task (#1370) 2024-09-10 01:17:37 -07:00
Liangsheng Yin 69b3bb9ae1 Unify forward mode (#1360) 2024-09-09 13:49:29 -07:00
Lianmin Zheng e4d68afcf0 [Minor] Many cleanup (#1357) 2024-09-09 04:14:11 -07:00
Kaichen Zhang - NTU 662ecd9368 [Feat] Add modalities for vision server when handling pixel values for llava (#1346) 2024-09-09 02:07:34 -07:00
Liangsheng Yin 05bea6883c Fix some online scheduling delay (#1345) 2024-09-07 20:46:27 -07:00
Liangsheng Yin ab4a83b259 Optimize schedule (#1339) 2024-09-05 14:30:26 -07:00
Lianmin Zheng 1e495e0847 [Fix] Fix select by ensuring each request has at least one token (#1318) 2024-09-03 06:31:45 -07:00
Lianmin Zheng f64eae3a29 [Fix] Reduce memory usage for loading llava model & Remove EntryClassRemapping (#1308) 2024-09-02 21:44:45 -07:00
Kai-Hsun ChenandYineng Zhang 0836055324 [Chore] Rename model_overide_args to model_override_args (#1284)
Signed-off-by: Kai-Hsun Chen <kaihsun@anyscale.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2024-09-01 03:14:56 -07:00
Liangsheng Yin 381dd57bd6 Sampler cudagraph (#1253) 2024-08-28 18:58:52 -07:00
Zhiqiang Xie 8153168c96 fix data racing due to mutable reference using deepcopy (#1255) 2024-08-28 18:57:54 -07:00
Lianmin Zheng bf53bf5142 [Fix] Fix llava on multi images (#1247) 2024-08-28 06:33:05 -07:00
Yineng Zhang f25f4dfde5 hotfix: revert sampler CUDA Graph (#1242) 2024-08-28 21:16:47 +10:00
havetc 909f34363b [FIX] Wrong logger (#1230) 2024-08-27 20:10:46 +10:00
havetcandYineng Zhang 9935f97b3e [FEAT] JSON constrained support (#1125)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2024-08-26 09:37:26 -07:00
Liangsheng YinandYineng Zhang 75ce37f401 Move sampler into CUDA graph (#1201)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2024-08-26 07:02:50 -07:00
Liangsheng Yin 632d506d0b minor: improve CI and dependencies (#1212) 2024-08-26 04:26:31 +00:00
Kaichen Zhang - NTU 3579162ab1 [Fix] Multi-images loading error (#1218) 2024-08-26 03:58:51 +00:00
Lianmin Zheng 902278008a [Minor] Improve the function organization in TokenizerManager & improve loggers (#1208) 2024-08-25 14:46:34 -07:00
ChayenneandYing Sheng 30b4f771b0 Support Alibaba-NLP/gte-Qwen2-7B-instruct embedding Model (#1186)
Co-authored-by: Ying Sheng <sqy1415@gmail.com>
2024-08-25 10:29:12 -07:00
Kaichen Zhang - NTU 66e7dcaf70 [Fix] Fixing the multi-images error for llava-onevision (#1205) 2024-08-25 10:28:23 -07:00
Ying Sheng 1cb4da5c5f [Fix] the issue of random order when input is a list (#1199) 2024-08-24 21:43:03 -07:00
Lianmin Zheng f6af3a6561 Cleanup readme, llava examples, usage examples and nccl init (#1194) 2024-08-24 08:02:23 -07:00
Kaichen Zhang - NTUandBo Li a5b14ad043 [Feat/WIP] add llava-onevision, with support for (1) siglip encoder, (2) qwen2 decoder (3) openai api compatible server. (#1123)
Co-authored-by: Bo Li <drluodian@gmail.com>
2024-08-23 14:11:16 -07:00
Liangsheng Yin 364d3d72a7 Fix broken penalty (#1184) 2024-08-22 08:16:35 +00:00
Lianmin Zheng 5623826f73 [Minor] Improve logging and rename the health check endpoint name (#1180) 2024-08-21 19:24:36 -07:00
Liangsheng Yin 83e23c69b3 Improve code style of sampler (#1168) 2024-08-21 16:48:24 -07:00
intervitens 068e9eae55 Support min-p sampling (#1167) 2024-08-21 22:49:32 +00:00