Commit Graph
100 Commits
Author SHA1 Message Date
Ying Shengandsglang-bot 15bc1f5cd7 Update .github/MAINTAINER.md (#13398)
Co-authored-by: sglang-bot <sglangbot@gmail.com>
2025-11-16 21:32:24 -08:00
168033d5fb Support mxfp4 for GPT-OSS (#8843)
Co-authored-by: Co-author fzyzcjy <ch271828n@outlook.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Co-authored-by: zhuofan1123 <zhuofanl@nvidia.com>
Co-authored-by: liz-badada <jinyanc@nvidia.com>
Co-authored-by: xutizhou <xutingz@nvidia.com>
Co-authored-by: linhu-nv <linhu@nvidia.com>
2025-08-06 00:05:25 -07:00
Ying Sheng c1d2061f97 Add initial support for gpt-oss (#8824) 2025-08-05 13:42:01 -07:00
Ying Sheng 42fc44100a [minor] Add server_args check for Llama4 with hybrid (#7988) 2025-07-12 20:13:40 -07:00
Ying Sheng ccfa084125 [script] update loogle test (#7975) 2025-07-12 00:06:17 -07:00
Ying Sheng bcc5ba94b4 [minor fix] SWA missing methods (#7972) 2025-07-11 23:57:02 -07:00
Ying Sheng cee9f329c4 [minor fix] llama4 hybrid memory (#7950) 2025-07-11 23:11:36 -07:00
Ying Sheng fb71725c98 Fix a bug in schedule_policy (#6276) 2025-05-13 18:04:00 -07:00
Ying Sheng bad7c26fdc [PP] Fix init_memory_pool desync & add PP for mixtral (#6223) 2025-05-12 12:38:09 -07:00
Ying Sheng 11383cec3c [PP] Add pipeline parallelism (#5724) 2025-04-30 18:18:07 -07:00
Ying Sheng d7bc19a46a add multi-lora feature in README.md (#5463) 2025-04-16 03:25:25 -07:00
Ying ShengandSehoon Kim 1b859295f4 [Eagle] Remove the greedy branch and some redundant code (#4363)
Co-authored-by: Sehoon Kim <sehoon@x.ai>
2025-03-16 02:48:55 -07:00
Ying Sheng 52a34d7448 Add greedy verification kernel (#4383) 2025-03-16 00:58:26 -07:00
Ying Sheng 34c8898755 Check eagle server args (#4217) 2025-03-09 01:10:43 -08:00
02e9e9f1cf Add codeowners for eagle implementations (#4131)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: Sehoon Kim <kssteven418@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2025-03-05 23:16:49 -08:00
Ying ShengandKe Bao d3d4d76758 [Eagle] Refactor eagle speculative decoding (#3986)
Co-authored-by: Ke Bao <ISPObaoke@163.com>
2025-03-05 08:06:07 -08:00
Ying Sheng d23cb9a01e [Eagle] reduce one draft forward (#3468) 2025-02-10 20:21:49 +08:00
Ying Sheng 52a492a16e Update contribution_guide.md (#3452) 2025-02-10 12:53:47 +08:00
Ying Sheng 7b4e61fff3 [Fix] Fix eagle with disable cuda graph (#3411) 2025-02-09 08:40:00 +08:00
Ying Sheng cde4bbd5cc docs: add Novita for adoption and sponsorship (#3227) 2025-01-30 18:28:22 -08:00
Ying Sheng dc7eb01f19 [Fix] fix openai adapter (#2685) 2024-12-31 10:48:19 +00:00
Ying Sheng e0e09fceeb [Session] Update session control interface (#2635) 2024-12-29 02:10:27 -08:00
Ying Sheng 8a56b43175 [Bench] Flush cache before benchmarking (#2566) 2024-12-24 11:21:21 +08:00
Ying Sheng 8586b72da0 [feat] Enable chunked prefill for llava-onevision (#2412) 2024-12-09 09:52:38 -08:00
Ying Sheng aa47f64223 Revert "[feat] Enable chunked prefill for llava-onevision" (#2329) 2024-12-02 23:11:13 -08:00
Ying Sheng 480e38a733 [feat] Enable chunked prefill for llava-onevision (#2281) 2024-12-02 20:19:02 -08:00
8b48496aaf Revert "Revert "Add simple CPU offloading support"" (#2253)
Co-authored-by: Jani Monoses <jani.monoses@gmail.com>
Co-authored-by: youkaichao <youkaichao@gmail.com>
2024-11-28 23:58:54 -08:00
Ying Sheng 4057ea82c9 Revert "Add simple CPU offloading support" (#2252)
We'll re-add the commit to correctly ack Kaichao's authorship
2024-11-28 23:36:55 -08:00
Ying Sheng b7038fec9b [fix] Fix prefix caching for multi-image/video (#2239) 2024-11-28 12:08:13 -08:00
Ying Sheng 37c8a5761f [feat] Support session control for vision language models (#2210) 2024-11-27 00:03:29 -08:00
Ying Sheng e1e595d702 [feat] Refactor session control interface and add CI (#2173) 2024-11-25 12:32:51 -08:00
Ying Sheng 5942dfc00a [feat] Add session control (#2073) 2024-11-20 00:36:53 -08:00
Ying Sheng 4e2af03cfa [Production] Drain requests before exit when receive SIGTERM (#1838) 2024-10-30 10:22:56 -07:00
Ying ShengandByron Hsu 2fce449b1c [API] add get memory pool size (#1760)
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
2024-10-23 07:02:29 +00:00
Ying Sheng 95946271af Update README.md 2024-10-19 22:29:12 -07:00
Ying Sheng 5c4ce65631 Update README.md (#1722) 2024-10-19 22:27:38 -07:00
Ying Sheng e4b367baa8 [Event] Add online meetup meeting link (#1686) 2024-10-16 10:58:14 -07:00
Ying Sheng 2725f8da61 [Minor] Rename no_eos_trim to no_stop_trim (#1661) 2024-10-13 20:30:03 -07:00
Ying Sheng 4876117171 [Fix] fix eos trim inconsistency (#1650) 2024-10-13 01:07:09 -07:00
Ying Sheng c5325aba75 [Profile] Add pytorch profiler (#1604) 2024-10-07 14:37:16 -07:00
Ying Sheng c98e84c21e [Minor, Performance] Use torch.argmax for greedy sampling (#1589) 2024-10-06 13:15:05 -07:00
Ying Sheng 9c064bf78a [LoRA, Performance] Speedup multi-LoRA serving - Step 1 (#1587) 2024-10-06 10:33:44 -07:00
Ying Sheng 1c1bdc7699 [Event] Update README.md (#1572) 2024-10-05 11:16:47 -07:00
Ying Shengandhnyls2002 04b262cd91 [Fix] Fix major performance bug in certain cases (#1563)
Co-authored-by: hnyls2002 <hnyls2002@gmail.com>
2024-10-04 08:51:11 +00:00
Ying Sheng f202ed9712 [Refactor] Simplify io_struct and tokenizer_manager (#1549) 2024-10-01 10:25:32 -07:00
Ying Sheng 0f4fb19bc8 [Fix, LoRA] fix LoRA with updates in main (#1545) 2024-09-30 10:06:08 -07:00
Ying Sheng 9aa6553d2a [Feature] Support reward model LxzGordon/URM-LLaMa-3.1-8B (#1525) 2024-09-27 23:32:11 -07:00
Ying Sheng b1e330bcb0 [Event] Update meeting link (#1529) 2024-09-27 13:30:04 -07:00
Ying Sheng 37c5899fc2 Release v0.3.2 (#1512) 2024-09-25 14:17:09 +08:00
Ying Sheng f39a0197fd Revert "kernel: use tensor cores for flashinfer gqa kernels" (#1511) 2024-09-24 22:50:31 -07:00
Ying Sheng e4780cf839 [API, Feature] Support response prefill for openai API (#1490) 2024-09-22 06:46:17 -07:00
Ying Sheng 6f3cf1297e [CI, AMD] Add AMD tests to CI (#1491) 2024-09-22 04:45:10 -07:00
Ying Sheng 8f527e2940 [Event] Add public meeting invite to README (#1458) 2024-09-18 23:53:22 +08:00
Ying Sheng 2abe4f1cb6 Revert "[Minor] Raise exception for wrong import (#1409)" (#1432) 2024-09-15 15:22:32 -07:00
Ying Sheng 37963394aa [Feature] Support LoRA path renaming and add LoRA serving benchmarks (#1433) 2024-09-15 12:46:04 -07:00
Ying Sheng 9a903a8784 [Minor] Raise exception for wrong import (#1409) 2024-09-12 23:02:36 -07:00
Ying Sheng eb02c1618a [Minor, CI] remove lora test from minimal suite (#1406) 2024-09-12 16:49:50 -07:00
Ying Sheng 712216928f [Feature] Initial support for multi-LoRA serving (#1307) 2024-09-12 16:46:14 -07:00
Ying Sheng 689ff588ec [CI] Return output logprobs in unit test (#1361) 2024-09-09 13:05:13 -07:00
Ying Sheng 308d024092 [CI] Fix the issue of unit test hanging (#1211) 2024-08-25 16:21:37 -07:00
Ying Sheng ab4990e4bf [Minor] Temporarily skip flaky test (#1209) 2024-08-25 14:49:23 -07:00
Ying Sheng 1cb4da5c5f [Fix] the issue of random order when input is a list (#1199) 2024-08-24 21:43:03 -07:00
Ying Sheng e61d13acdf [CI] Fix the problem of hf runner too slow (#1202) 2024-08-24 18:35:55 -07:00
Ying Sheng 5fafcac008 Fix benchmark script (#1185) 2024-08-22 09:03:25 +00:00
Ying Sheng 93d4e354d8 [Fix] Window attention compatible with RadixAttention and chunked prefill (#1112) 2024-08-15 10:33:20 -07:00
Ying Sheng 14cb544d56 [Fix] fix flashinfer usage for window attention (#1107) 2024-08-15 00:53:24 -07:00
Ying Sheng 8d2d876fc8 [Fix] fix the typo bug for window attention (#1106) 2024-08-14 21:56:01 -07:00
Ying Sheng 6767e2229f Support jinja as chat template file (#1104) 2024-08-14 17:43:14 -07:00
Ying Sheng 96a2093ef0 [Fix] Compatibility of window attention and cuda graph (#1090) 2024-08-14 10:37:01 -07:00
Ying Sheng 0909bb0d2f [Feat] Add window attention for gemma-2 (#1056) 2024-08-13 17:01:26 -07:00
Ying Sheng 32f6144323 fix: Fix returned prefill logits and add output str test (#1046) 2024-08-12 06:13:45 +00:00
Ying Sheng b68c4c073b fix: force max new tokens to be 1 for embedding request (#1019) 2024-08-10 13:46:42 -07:00
Ying Sheng 7599badeaf Support embedding input as a list (#1014) 2024-08-10 08:39:05 -07:00
Ying Sheng b16e856f11 Add openai embedding API (#997) 2024-08-09 11:19:18 -07:00
Ying Sheng e040a2450b Add e5-mistral embedding model - step 3/3 (#988) 2024-08-08 16:31:19 -07:00
Ying Sheng 9f662501a3 Move torch.compile configs into cuda_graph_runner.py (#993) 2024-08-08 13:20:30 -07:00
Ying Sheng 228cf47547 Create contributor_guide.md (#992) 2024-08-08 03:58:47 -07:00
Ying Sheng 20a4f927dc Add io struct for embedding models [unreachable code] - step 2/3 (#987) 2024-08-08 07:52:31 +00:00
Ying Sheng 0de7c2d09e Add e5-mistral modules [unreachable code] - step 1/3 (#983) 2024-08-08 00:04:15 -07:00
Ying Sheng 00023d622a [minor] Update type annotation in tokenizer_manager.py (#982) 2024-08-08 01:48:45 +00:00
Ying Sheng ff68ae857a Show more error messages for warmup errors (#932) 2024-08-06 23:57:06 -07:00
Ying Sheng 399cad91f3 Update README.md (#927) 2024-08-04 23:01:35 -07:00
Ying Sheng 0a4f5f9bea Test regex in vision api (#926) 2024-08-04 22:52:41 -07:00
Ying Sheng 3bc99e6fe4 Test openai vision api (#925) 2024-08-05 13:51:55 +10:00
Ying Sheng 141e8c71a3 Bump version to 0.2.10 (#923) 2024-08-04 16:52:51 -07:00
Ying Sheng 975adb802b Update hyperparameter_tuning.md (#918) 2024-08-04 13:51:52 -07:00
Ying Sheng 0d4f3a9fcd Make API Key OpenAI-compatible (#917) 2024-08-04 13:35:44 -07:00
Ying Sheng 995af5a54b Improve the structure of CI (#911) 2024-08-03 23:09:21 -07:00
Ying Sheng 70cc0749ce Add model accuracy test - step 1 (#866) 2024-08-03 18:20:50 -07:00
Ying Sheng 8c5382e62c Update README.md 2024-08-03 12:58:41 -07:00
Ying Sheng 001b0bdd08 Update the base image of the docker (#900) 2024-08-02 21:54:57 -07:00
Ying Sheng b906c01592 Bump version to 0.2.9.post1 (#899) 2024-08-02 12:08:00 -07:00
Ying Sheng 30a9b2ef20 Bump version to v0.2.9 (#890) 2024-08-02 01:45:48 -07:00
Ying Sheng 3cadecf0c4 Increase openai client limit (#886) 2024-08-02 00:47:23 -07:00
Ying Sheng e90e3a50d4 Add benchmark: HumanEval (#889) 2024-08-02 00:46:41 -07:00
Ying Sheng fbd6b94d69 Fix the double BOS problem in the HF chat template (#888) 2024-08-02 00:30:50 -07:00
Ying Sheng 4c8093c8db Update workflow name (#883) 2024-08-01 21:29:46 -07:00
Ying Sheng ae7ee01a8e Add accuracy test to CI: MMLU (#882) 2024-08-01 21:20:17 -07:00
Ying Sheng 76e59088d8 Add more unit tests to CI (#880) 2024-08-01 18:14:33 -07:00
Ying Sheng 60340a3643 Improve the coverage of the openai api server test (#878) 2024-08-01 16:01:30 -07:00