Commit Graph

  • 446ea33277 fix: creat new dict everytime for putting new frame (#1464) Li Bo 2024-09-19 16:31:48 +08:00
  • 8f527e2940 [Event] Add public meeting invite to README (#1458) Ying Sheng 2024-09-18 08:53:22 -07:00
  • 7f24ea95c3 Fuse top_k and top_k in the sampler (#1457) Lianmin Zheng 2024-09-18 04:35:35 -07:00
  • 1acccb364a Fix oom issues with fp8 for llama (#1454) Lianmin Zheng 2024-09-18 03:45:19 -07:00
  • aa2750beb3 [Bugfix] Enable SGLang on AMD GPUs via PyTorch for ROCm (#1419) (#1453) HAI 2024-09-18 02:01:35 -07:00
  • 5e62a6b706 Add bench_server_latency.py (#1452) Lianmin Zheng 2024-09-18 00:56:06 -07:00
  • 5752f25eef Fixed n>1 causing list index out of range with VLM (#1449) Xiao Yu 2024-09-18 03:46:32 -04:00
  • 7c162fa9c5 Fix schedule bug (#1451) Liangsheng Yin 2024-09-17 22:59:32 -07:00
  • 36078fb247 fix schedule bug (#1450) Liangsheng Yin 2024-09-17 16:33:53 -07:00
  • b3710d2c93 Fix attention backend (#1448) Ke Bao 2024-09-17 22:07:53 +08:00
  • c6b6d2e71b Enable MLA by default (#1447) Ke Bao 2024-09-17 19:42:48 +08:00
  • 90a26be31c Release 0.3.1.post1 (#1445) Lianmin Zheng 2024-09-17 01:47:31 -07:00
  • 1f4b5f770d Add OLMoE model (#1444) Jani Monoses 2024-09-17 11:14:53 +03:00
  • 76524b70d1 Fix torch compile for deepseek-v2 (#1442) Ke Bao 2024-09-17 15:52:08 +08:00
  • 3a6e04185b [Feature, Hardware] Enable SGLang on AMD GPUs via PyTorch for ROCm (#1420) HAI 2024-09-17 00:43:52 -07:00
  • 2fa5cec775 Simplify sampler and its error handling (#1441) Lianmin Zheng 2024-09-16 21:23:31 -07:00
  • 27b557aea7 Clean up model loader (#1440) Lianmin Zheng 2024-09-16 18:16:27 -07:00
  • 93dffd699b Add constrained_json_whitespace_pattern to ServerArgs (#1438) zifeitong 2024-09-16 13:29:18 -07:00
  • 2abe4f1cb6 Revert "[Minor] Raise exception for wrong import (#1409)" (#1432) Ying Sheng 2024-09-15 15:22:32 -07:00
  • 37963394aa [Feature] Support LoRA path renaming and add LoRA serving benchmarks (#1433) Ying Sheng 2024-09-15 12:46:04 -07:00
  • 899cf5c438 Remove deprecated configs (#1431) Lianmin Zheng 2024-09-15 08:52:18 -07:00
  • e79f6cd73d Release v0.3.1 (#1430) Lianmin Zheng 2024-09-15 07:03:16 -07:00
  • 9ba1f09760 [Fix] Fix logprob and normalized_logprob (#1428) Lianmin Zheng 2024-09-15 06:36:06 -07:00
  • 282681b8a1 Update backend.md (#1429) Lianmin Zheng 2024-09-15 02:55:34 -07:00
  • 58cafe23a7 Add libibverbs-dev to Dockerfile (#1427) William Arnold 2024-09-15 15:40:31 +09:00
  • 9463bc1385 Enable torch.compile for triton backend (#1422) Lianmin Zheng 2024-09-14 15:38:37 -07:00
  • e3fc4658f4 fix: resolve nightly eval (#1426) Yineng Zhang 2024-09-15 01:07:52 +09:00
  • 33b54e7c40 Add pytorch sampling backend ut (#1425) Ke Bao 2024-09-14 23:15:30 +08:00
  • 30b404ce72 Add torchao quant for mixtral and qwen_moe (#1418) Jerry Zhang 2024-09-13 23:46:55 -07:00
  • 70b6802982 Optimize conflicts between CUDA graph and vocab mask tensors (#1392) Liangsheng Yin 2024-09-13 20:27:53 -07:00
  • f3d32f888a ci: fix finish (#1414) Yineng Zhang 2024-09-14 00:01:30 +09:00
  • 8779da95d6 Update pr-test.yml (#1412) Lianmin Zheng 2024-09-13 00:37:13 -07:00
  • ad0ff62a4c Balance test in CI (#1411) Lianmin Zheng 2024-09-12 23:29:44 -07:00
  • 9a903a8784 [Minor] Raise exception for wrong import (#1409) Ying Sheng 2024-09-12 23:02:36 -07:00
  • 68be2f6d3b [CI] Include triton backend and online serving benchmark into CI (#1408) Lianmin Zheng 2024-09-12 21:36:41 -07:00
  • b912de11b0 Make stop reason a dict instead of str (#1407) Lianmin Zheng 2024-09-12 20:47:31 -07:00
  • eb02c1618a [Minor, CI] remove lora test from minimal suite (#1406) Ying Sheng 2024-09-12 16:49:50 -07:00
  • 712216928f [Feature] Initial support for multi-LoRA serving (#1307) Ying Sheng 2024-09-12 16:46:14 -07:00
  • c33d82a211 Add Support for XVERSE Models (Dense and MoE) to sglang (#1397) hxer7963 2024-09-12 16:47:52 +08:00
  • 8234e663e9 [Minor Fix] Fix llava modalities issue for single-image (#1402) Kaichen Zhang - NTU 2024-09-12 16:10:26 +08:00
  • debbdb5178 kernel: use tensor cores for flashinfer gqa kernels (#1403) Zihao Ye 2024-09-12 00:38:18 -07:00
  • 3efa798116 Support cuda graph in the triton attention backend (#1401) Lianmin Zheng 2024-09-12 00:36:55 -07:00
  • 2a71be5e25 Fix README format (#1399) William 2024-09-12 14:46:51 +08:00
  • 4462137777 Add no commit to main rule (#1393) Liangsheng Yin 2024-09-11 14:40:45 -07:00
  • fec185ce0c Refactor attention backend (#1381) Lianmin Zheng 2024-09-11 11:44:26 -07:00
  • c03cece42f Improve error reporting during server launch (#1390) Lianmin Zheng 2024-09-11 04:50:04 -07:00
  • 15c75e4146 [Fix] Fix --disable-flashinfer (#1389) Lianmin Zheng 2024-09-11 04:36:21 -07:00
  • 224200e3c2 BaiChuan2 Model (#1367) Vectory 2024-09-11 18:55:24 +08:00
  • 8c0efa514d remove assertion in triton attention and add an unit test (#1385) Byron Hsu 2024-09-11 03:22:07 -07:00
  • 144bc70fcc Organize flashinfer indices update (#1378) Liangsheng Yin 2024-09-10 17:38:59 -07:00
  • 46094e0c1b Deprecate --disable-flashinfer and introduce --attention-backend (#1380) Lianmin Zheng 2024-09-10 17:11:16 -07:00
  • 3a6e8b6d78 [Minor] move triton attention kernels into a separate folder (#1379) Lianmin Zheng 2024-09-10 15:15:08 -07:00
  • fbb4754cb8 Fix vocab mask update bug (#1376) Liangsheng Yin 2024-09-10 13:10:36 -07:00
  • 6c7cb90365 [Minor] improve kill scripts and torchao import (#1375) Lianmin Zheng 2024-09-10 11:27:03 -07:00
  • dff2860a69 Fix CORS compatibility with OpenAI, vLLM, TGI, LMDeploy (#1373) josephrocca 2024-09-11 00:35:03 +08:00
  • e72275cf7f Support MiniCPM3 (#1371) William 2024-09-10 17:57:52 +08:00
  • fec2d1223c [Fix] fix bug of undefined is_single in meth create_abort_task (#1370) wangchao 2024-09-10 16:17:37 +08:00
  • 8d1095dbf0 [Docs] Improve documentations (#1368) Lianmin Zheng 2024-09-09 20:48:28 -07:00
  • 743007e1ce Adding Documentation for installation (#1300) Chayenne 2024-09-10 10:09:13 +08:00
  • 9144ed1067 Support OpenAI API json_schema response format (#1363) zifeitong 2024-09-09 19:08:25 -07:00
  • 69b3bb9ae1 Unify forward mode (#1360) Liangsheng Yin 2024-09-09 13:49:29 -07:00
  • 689ff588ec [CI] Return output logprobs in unit test (#1361) Ying Sheng 2024-09-09 13:05:13 -07:00
  • a7c47e0f02 Add torchao quant (int4/int8/fp8) to llama models (#1341) Jerry Zhang 2024-09-09 05:32:41 -07:00
  • e4d68afcf0 [Minor] Many cleanup (#1357) Lianmin Zheng 2024-09-09 04:14:11 -07:00
  • c9b75917d5 [server] Passing model_override_args to launch_server via the CLI. (#1298) Kai-Hsun Chen 2024-09-09 02:14:25 -07:00
  • 662ecd9368 [Feat] Add modalities for vision server when handling pixel values for llava (#1346) Kaichen Zhang - NTU 2024-09-09 17:07:34 +08:00
  • 8e6bdf851c [triton] Support head_dim not 2^n in triton extend and decode attention (#1281) Byron Hsu 2024-09-09 01:30:24 -07:00
  • 05bea6883c Fix some online scheduling delay (#1345) Liangsheng Yin 2024-09-07 20:46:27 -07:00
  • ab4a83b259 Optimize schedule (#1339) Liangsheng Yin 2024-09-05 14:30:26 -07:00
  • 62f15eea5a docs: add conclusion (#1340) Yineng Zhang 2024-09-06 04:25:14 +10:00
  • 79794af52d docs: highlight ttft itl and throughput (#1337) Yineng Zhang 2024-09-06 00:00:06 +10:00
  • 3494b32c3a docs: update README (#1336) Yineng Zhang 2024-09-05 23:39:44 +10:00
  • eda7c09048 Remove useless fields in global_config.py (#1328) Lianmin Zheng 2024-09-04 05:37:32 -07:00
  • 5ab9418f5b [Doc] update news (#1327) Yineng Zhang 2024-09-04 21:21:21 +10:00
  • 843e63d809 Fix the flaky test test_moe_eval_accuracy_large.py (#1326) Lianmin Zheng 2024-09-04 04:15:11 -07:00
  • a63c8275c6 chore: bump v0.3.0 (#1320) Yineng Zhang 2024-09-04 06:32:18 +10:00
  • dc67d97693 misc: speedup load safetensors (#1319) Yineng Zhang 2024-09-04 04:29:53 +10:00
  • 1e495e0847 [Fix] Fix select by ensuring each request has at least one token (#1318) Lianmin Zheng 2024-09-03 06:31:45 -07:00
  • 12cb115d38 Fix llama2 weight loader (#1317) Lianmin Zheng 2024-09-03 05:32:14 -07:00
  • c500f96bb1 Update README.md for llava-onevision instructions (#1313) Lianmin Zheng 2024-09-03 01:43:08 -07:00
  • 474317f2b6 Support Phi3 mini and medium (#1299) Jani Monoses 2024-09-03 07:49:40 +03:00
  • f64eae3a29 [Fix] Reduce memory usage for loading llava model & Remove EntryClassRemapping (#1308) Lianmin Zheng 2024-09-02 21:44:45 -07:00
  • a5a134f39f Fix bugs in sampler with CUDA graph / torch.compile (#1306) Liangsheng Yin 2024-09-02 16:18:48 -07:00
  • 2561ed012c feat: update nightly gsm8k eval (#1304) Yineng Zhang 2024-09-03 01:18:41 +10:00
  • 9999442756 Release v0.2.15 (#1295) Lianmin Zheng 2024-09-01 22:22:38 -07:00
  • 6def9b018c Fix hang when doing s += None. (#1297) Max Shawabkeh 2024-09-01 21:56:33 -07:00
  • 47f20da223 Fix regex mask (#1296) Liangsheng Yin 2024-09-01 21:50:58 -07:00
  • 4a9f8ea43b [doc] Fix more broken links (#1294) Byron Hsu 2024-09-01 14:46:36 -07:00
  • 58fa607622 Fix the flaky tests in test_moe_eval_accuracy_large.py (#1293) Lianmin Zheng 2024-09-01 12:20:46 -07:00
  • 6487ef64c6 ci: add nightly eval (#1291) Yineng Zhang 2024-09-02 03:19:49 +10:00
  • 9b0805242e fix: resolve fp8 for mixtral (#1290) Yineng Zhang 2024-09-02 00:29:06 +10:00
  • 32a4141d5a Allow new lines during JSON generation (#1277) Enrique Shockwave 2024-09-01 11:42:29 +01:00
  • 0836055324 [Chore] Rename model_overide_args to model_override_args (#1284) Kai-Hsun Chen 2024-09-01 03:14:56 -07:00
  • 00b19f198f [triton] Remove the zero initialization of qk_acc by directly writing the result (#1288) Byron Hsu 2024-09-01 03:12:06 -07:00
  • 6cb32ef92c Support Triton fp8 e5m2 kv cache (#1286) Ke Bao 2024-09-01 17:46:40 +08:00
  • 761b2cebd6 [CI] merge all ci tests into one file (#1289) Lianmin Zheng 2024-09-01 02:36:56 -07:00
  • 54772f784a feat: fix fp8 for MLA and support bmm fp8 for DeepSeek V2 (#1285) Yineng Zhang 2024-09-01 17:28:06 +10:00
  • 1b5d56f7f8 [CI] Add more multi-gpu tests (#1280) Lianmin Zheng 2024-09-01 00:27:25 -07:00
  • d134c139a1 Optimize the update flashinfer indices (#1262) xiaobochen 2024-08-31 23:40:28 -07:00
  • 6cc9c52521 [doc] fix quick start link (#1282) Byron Hsu 2024-08-31 22:54:34 -07:00