 Ying Shengandsglang-bot
|
15bc1f5cd7
|
Update .github/MAINTAINER.md (#13398)
Co-authored-by: sglang-bot <sglangbot@gmail.com>
|
2025-11-16 21:32:24 -08:00 |
|
     
|
168033d5fb
|
Support mxfp4 for GPT-OSS (#8843)
Co-authored-by: Co-author fzyzcjy <ch271828n@outlook.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Co-authored-by: zhuofan1123 <zhuofanl@nvidia.com>
Co-authored-by: liz-badada <jinyanc@nvidia.com>
Co-authored-by: xutizhou <xutingz@nvidia.com>
Co-authored-by: linhu-nv <linhu@nvidia.com>
|
2025-08-06 00:05:25 -07:00 |
|
Ying Sheng
|
c1d2061f97
|
Add initial support for gpt-oss (#8824)
|
2025-08-05 13:42:01 -07:00 |
|
Ying Sheng
|
42fc44100a
|
[minor] Add server_args check for Llama4 with hybrid (#7988)
|
2025-07-12 20:13:40 -07:00 |
|
Ying Sheng
|
ccfa084125
|
[script] update loogle test (#7975)
|
2025-07-12 00:06:17 -07:00 |
|
Ying Sheng
|
bcc5ba94b4
|
[minor fix] SWA missing methods (#7972)
|
2025-07-11 23:57:02 -07:00 |
|
Ying Sheng
|
cee9f329c4
|
[minor fix] llama4 hybrid memory (#7950)
|
2025-07-11 23:11:36 -07:00 |
|
Ying Sheng
|
fb71725c98
|
Fix a bug in schedule_policy (#6276)
|
2025-05-13 18:04:00 -07:00 |
|
Ying Sheng
|
bad7c26fdc
|
[PP] Fix init_memory_pool desync & add PP for mixtral (#6223)
|
2025-05-12 12:38:09 -07:00 |
|
Ying Sheng
|
11383cec3c
|
[PP] Add pipeline parallelism (#5724)
|
2025-04-30 18:18:07 -07:00 |
|
Ying Sheng
|
d7bc19a46a
|
add multi-lora feature in README.md (#5463)
|
2025-04-16 03:25:25 -07:00 |
|
 Ying ShengandSehoon Kim
|
1b859295f4
|
[Eagle] Remove the greedy branch and some redundant code (#4363)
Co-authored-by: Sehoon Kim <sehoon@x.ai>
|
2025-03-16 02:48:55 -07:00 |
|
Ying Sheng
|
52a34d7448
|
Add greedy verification kernel (#4383)
|
2025-03-16 00:58:26 -07:00 |
|
Ying Sheng
|
34c8898755
|
Check eagle server args (#4217)
|
2025-03-09 01:10:43 -08:00 |
|
  
|
02e9e9f1cf
|
Add codeowners for eagle implementations (#4131)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: Sehoon Kim <kssteven418@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2025-03-05 23:16:49 -08:00 |
|
 Ying ShengandKe Bao
|
d3d4d76758
|
[Eagle] Refactor eagle speculative decoding (#3986)
Co-authored-by: Ke Bao <ISPObaoke@163.com>
|
2025-03-05 08:06:07 -08:00 |
|
Ying Sheng
|
d23cb9a01e
|
[Eagle] reduce one draft forward (#3468)
|
2025-02-10 20:21:49 +08:00 |
|
Ying Sheng
|
52a492a16e
|
Update contribution_guide.md (#3452)
|
2025-02-10 12:53:47 +08:00 |
|
Ying Sheng
|
7b4e61fff3
|
[Fix] Fix eagle with disable cuda graph (#3411)
|
2025-02-09 08:40:00 +08:00 |
|
Ying Sheng
|
cde4bbd5cc
|
docs: add Novita for adoption and sponsorship (#3227)
|
2025-01-30 18:28:22 -08:00 |
|
Ying Sheng
|
dc7eb01f19
|
[Fix] fix openai adapter (#2685)
|
2024-12-31 10:48:19 +00:00 |
|
Ying Sheng
|
e0e09fceeb
|
[Session] Update session control interface (#2635)
|
2024-12-29 02:10:27 -08:00 |
|
Ying Sheng
|
8a56b43175
|
[Bench] Flush cache before benchmarking (#2566)
|
2024-12-24 11:21:21 +08:00 |
|
Ying Sheng
|
8586b72da0
|
[feat] Enable chunked prefill for llava-onevision (#2412)
|
2024-12-09 09:52:38 -08:00 |
|
Ying Sheng
|
aa47f64223
|
Revert "[feat] Enable chunked prefill for llava-onevision" (#2329)
|
2024-12-02 23:11:13 -08:00 |
|
Ying Sheng
|
480e38a733
|
[feat] Enable chunked prefill for llava-onevision (#2281)
|
2024-12-02 20:19:02 -08:00 |
|
 
|
8b48496aaf
|
Revert "Revert "Add simple CPU offloading support"" (#2253)
Co-authored-by: Jani Monoses <jani.monoses@gmail.com>
Co-authored-by: youkaichao <youkaichao@gmail.com>
|
2024-11-28 23:58:54 -08:00 |
|
Ying Sheng
|
4057ea82c9
|
Revert "Add simple CPU offloading support" (#2252)
We'll re-add the commit to correctly ack Kaichao's authorship
|
2024-11-28 23:36:55 -08:00 |
|
Ying Sheng
|
b7038fec9b
|
[fix] Fix prefix caching for multi-image/video (#2239)
|
2024-11-28 12:08:13 -08:00 |
|
Ying Sheng
|
37c8a5761f
|
[feat] Support session control for vision language models (#2210)
|
2024-11-27 00:03:29 -08:00 |
|
Ying Sheng
|
e1e595d702
|
[feat] Refactor session control interface and add CI (#2173)
|
2024-11-25 12:32:51 -08:00 |
|
Ying Sheng
|
5942dfc00a
|
[feat] Add session control (#2073)
|
2024-11-20 00:36:53 -08:00 |
|
Ying Sheng
|
4e2af03cfa
|
[Production] Drain requests before exit when receive SIGTERM (#1838)
|
2024-10-30 10:22:56 -07:00 |
|
 Ying ShengandByron Hsu
|
2fce449b1c
|
[API] add get memory pool size (#1760)
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
|
2024-10-23 07:02:29 +00:00 |
|
Ying Sheng
|
95946271af
|
Update README.md
|
2024-10-19 22:29:12 -07:00 |
|
Ying Sheng
|
5c4ce65631
|
Update README.md (#1722)
|
2024-10-19 22:27:38 -07:00 |
|
Ying Sheng
|
e4b367baa8
|
[Event] Add online meetup meeting link (#1686)
|
2024-10-16 10:58:14 -07:00 |
|
Ying Sheng
|
2725f8da61
|
[Minor] Rename no_eos_trim to no_stop_trim (#1661)
|
2024-10-13 20:30:03 -07:00 |
|
Ying Sheng
|
4876117171
|
[Fix] fix eos trim inconsistency (#1650)
|
2024-10-13 01:07:09 -07:00 |
|
Ying Sheng
|
c5325aba75
|
[Profile] Add pytorch profiler (#1604)
|
2024-10-07 14:37:16 -07:00 |
|
Ying Sheng
|
c98e84c21e
|
[Minor, Performance] Use torch.argmax for greedy sampling (#1589)
|
2024-10-06 13:15:05 -07:00 |
|
Ying Sheng
|
9c064bf78a
|
[LoRA, Performance] Speedup multi-LoRA serving - Step 1 (#1587)
|
2024-10-06 10:33:44 -07:00 |
|
Ying Sheng
|
1c1bdc7699
|
[Event] Update README.md (#1572)
|
2024-10-05 11:16:47 -07:00 |
|
 Ying Shengandhnyls2002
|
04b262cd91
|
[Fix] Fix major performance bug in certain cases (#1563)
Co-authored-by: hnyls2002 <hnyls2002@gmail.com>
|
2024-10-04 08:51:11 +00:00 |
|
Ying Sheng
|
f202ed9712
|
[Refactor] Simplify io_struct and tokenizer_manager (#1549)
|
2024-10-01 10:25:32 -07:00 |
|
Ying Sheng
|
0f4fb19bc8
|
[Fix, LoRA] fix LoRA with updates in main (#1545)
|
2024-09-30 10:06:08 -07:00 |
|
Ying Sheng
|
9aa6553d2a
|
[Feature] Support reward model LxzGordon/URM-LLaMa-3.1-8B (#1525)
|
2024-09-27 23:32:11 -07:00 |
|
Ying Sheng
|
b1e330bcb0
|
[Event] Update meeting link (#1529)
|
2024-09-27 13:30:04 -07:00 |
|
Ying Sheng
|
37c5899fc2
|
Release v0.3.2 (#1512)
|
2024-09-25 14:17:09 +08:00 |
|
Ying Sheng
|
f39a0197fd
|
Revert "kernel: use tensor cores for flashinfer gqa kernels" (#1511)
|
2024-09-24 22:50:31 -07:00 |
|
Ying Sheng
|
e4780cf839
|
[API, Feature] Support response prefill for openai API (#1490)
|
2024-09-22 06:46:17 -07:00 |
|
Ying Sheng
|
6f3cf1297e
|
[CI, AMD] Add AMD tests to CI (#1491)
|
2024-09-22 04:45:10 -07:00 |
|
Ying Sheng
|
8f527e2940
|
[Event] Add public meeting invite to README (#1458)
|
2024-09-18 23:53:22 +08:00 |
|
Ying Sheng
|
2abe4f1cb6
|
Revert "[Minor] Raise exception for wrong import (#1409)" (#1432)
|
2024-09-15 15:22:32 -07:00 |
|
Ying Sheng
|
37963394aa
|
[Feature] Support LoRA path renaming and add LoRA serving benchmarks (#1433)
|
2024-09-15 12:46:04 -07:00 |
|
Ying Sheng
|
9a903a8784
|
[Minor] Raise exception for wrong import (#1409)
|
2024-09-12 23:02:36 -07:00 |
|
Ying Sheng
|
eb02c1618a
|
[Minor, CI] remove lora test from minimal suite (#1406)
|
2024-09-12 16:49:50 -07:00 |
|
Ying Sheng
|
712216928f
|
[Feature] Initial support for multi-LoRA serving (#1307)
|
2024-09-12 16:46:14 -07:00 |
|
Ying Sheng
|
689ff588ec
|
[CI] Return output logprobs in unit test (#1361)
|
2024-09-09 13:05:13 -07:00 |
|
Ying Sheng
|
308d024092
|
[CI] Fix the issue of unit test hanging (#1211)
|
2024-08-25 16:21:37 -07:00 |
|
Ying Sheng
|
ab4990e4bf
|
[Minor] Temporarily skip flaky test (#1209)
|
2024-08-25 14:49:23 -07:00 |
|
Ying Sheng
|
1cb4da5c5f
|
[Fix] the issue of random order when input is a list (#1199)
|
2024-08-24 21:43:03 -07:00 |
|
Ying Sheng
|
e61d13acdf
|
[CI] Fix the problem of hf runner too slow (#1202)
|
2024-08-24 18:35:55 -07:00 |
|
Ying Sheng
|
5fafcac008
|
Fix benchmark script (#1185)
|
2024-08-22 09:03:25 +00:00 |
|
Ying Sheng
|
93d4e354d8
|
[Fix] Window attention compatible with RadixAttention and chunked prefill (#1112)
|
2024-08-15 10:33:20 -07:00 |
|
Ying Sheng
|
14cb544d56
|
[Fix] fix flashinfer usage for window attention (#1107)
|
2024-08-15 00:53:24 -07:00 |
|
Ying Sheng
|
8d2d876fc8
|
[Fix] fix the typo bug for window attention (#1106)
|
2024-08-14 21:56:01 -07:00 |
|
Ying Sheng
|
6767e2229f
|
Support jinja as chat template file (#1104)
|
2024-08-14 17:43:14 -07:00 |
|
Ying Sheng
|
96a2093ef0
|
[Fix] Compatibility of window attention and cuda graph (#1090)
|
2024-08-14 10:37:01 -07:00 |
|
Ying Sheng
|
0909bb0d2f
|
[Feat] Add window attention for gemma-2 (#1056)
|
2024-08-13 17:01:26 -07:00 |
|
Ying Sheng
|
32f6144323
|
fix: Fix returned prefill logits and add output str test (#1046)
|
2024-08-12 06:13:45 +00:00 |
|
Ying Sheng
|
b68c4c073b
|
fix: force max new tokens to be 1 for embedding request (#1019)
|
2024-08-10 13:46:42 -07:00 |
|
Ying Sheng
|
7599badeaf
|
Support embedding input as a list (#1014)
|
2024-08-10 08:39:05 -07:00 |
|
Ying Sheng
|
b16e856f11
|
Add openai embedding API (#997)
|
2024-08-09 11:19:18 -07:00 |
|
Ying Sheng
|
e040a2450b
|
Add e5-mistral embedding model - step 3/3 (#988)
|
2024-08-08 16:31:19 -07:00 |
|
Ying Sheng
|
9f662501a3
|
Move torch.compile configs into cuda_graph_runner.py (#993)
|
2024-08-08 13:20:30 -07:00 |
|
Ying Sheng
|
228cf47547
|
Create contributor_guide.md (#992)
|
2024-08-08 03:58:47 -07:00 |
|
Ying Sheng
|
20a4f927dc
|
Add io struct for embedding models [unreachable code] - step 2/3 (#987)
|
2024-08-08 07:52:31 +00:00 |
|
Ying Sheng
|
0de7c2d09e
|
Add e5-mistral modules [unreachable code] - step 1/3 (#983)
|
2024-08-08 00:04:15 -07:00 |
|
Ying Sheng
|
00023d622a
|
[minor] Update type annotation in tokenizer_manager.py (#982)
|
2024-08-08 01:48:45 +00:00 |
|
Ying Sheng
|
ff68ae857a
|
Show more error messages for warmup errors (#932)
|
2024-08-06 23:57:06 -07:00 |
|
Ying Sheng
|
399cad91f3
|
Update README.md (#927)
|
2024-08-04 23:01:35 -07:00 |
|
Ying Sheng
|
0a4f5f9bea
|
Test regex in vision api (#926)
|
2024-08-04 22:52:41 -07:00 |
|
Ying Sheng
|
3bc99e6fe4
|
Test openai vision api (#925)
|
2024-08-05 13:51:55 +10:00 |
|
Ying Sheng
|
141e8c71a3
|
Bump version to 0.2.10 (#923)
|
2024-08-04 16:52:51 -07:00 |
|
Ying Sheng
|
975adb802b
|
Update hyperparameter_tuning.md (#918)
|
2024-08-04 13:51:52 -07:00 |
|
Ying Sheng
|
0d4f3a9fcd
|
Make API Key OpenAI-compatible (#917)
|
2024-08-04 13:35:44 -07:00 |
|
Ying Sheng
|
995af5a54b
|
Improve the structure of CI (#911)
|
2024-08-03 23:09:21 -07:00 |
|
Ying Sheng
|
70cc0749ce
|
Add model accuracy test - step 1 (#866)
|
2024-08-03 18:20:50 -07:00 |
|
Ying Sheng
|
8c5382e62c
|
Update README.md
|
2024-08-03 12:58:41 -07:00 |
|
Ying Sheng
|
001b0bdd08
|
Update the base image of the docker (#900)
|
2024-08-02 21:54:57 -07:00 |
|
Ying Sheng
|
b906c01592
|
Bump version to 0.2.9.post1 (#899)
|
2024-08-02 12:08:00 -07:00 |
|
Ying Sheng
|
30a9b2ef20
|
Bump version to v0.2.9 (#890)
|
2024-08-02 01:45:48 -07:00 |
|
Ying Sheng
|
3cadecf0c4
|
Increase openai client limit (#886)
|
2024-08-02 00:47:23 -07:00 |
|
Ying Sheng
|
e90e3a50d4
|
Add benchmark: HumanEval (#889)
|
2024-08-02 00:46:41 -07:00 |
|
Ying Sheng
|
fbd6b94d69
|
Fix the double BOS problem in the HF chat template (#888)
|
2024-08-02 00:30:50 -07:00 |
|
Ying Sheng
|
4c8093c8db
|
Update workflow name (#883)
|
2024-08-01 21:29:46 -07:00 |
|
Ying Sheng
|
ae7ee01a8e
|
Add accuracy test to CI: MMLU (#882)
|
2024-08-01 21:20:17 -07:00 |
|
Ying Sheng
|
76e59088d8
|
Add more unit tests to CI (#880)
|
2024-08-01 18:14:33 -07:00 |
|
Ying Sheng
|
60340a3643
|
Improve the coverage of the openai api server test (#878)
|
2024-08-01 16:01:30 -07:00 |
|