Zhiqiang Xie
|
f9eb04ddb2
|
upgrade sgl kernel to 0.2.1 for main (#7676)
|
2025-07-01 00:00:13 -07:00 |
|
Yineng Zhang
|
392e441ad1
|
chore: upgrade flashinfer v0.2.7 jit (#7663)
|
2025-06-30 13:26:26 -07:00 |
|
Lianmin Zheng
|
22352d47a9
|
Improve streaming, log_level, memory report, weight loading, and benchmark script (#7632)
Co-authored-by: Kan Wu <wukanustc@gmail.com>
|
2025-06-29 23:16:19 -07:00 |
|
Lianmin Zheng
|
071a1f51ae
|
[Minor] clean up multimodal processor and tokenizer manager (#7624)
|
2025-06-29 02:50:14 -07:00 |
|
Lifu Huang
|
49538d111b
|
Support dynamic LoRA loading / unloading in engine/server API (#7446)
|
2025-06-27 21:00:27 -07:00 |
|
mlmz
|
fe2a0f962f
|
minor: 'role' must be system/assistant/tool, but case insensitive for now (#7499)
|
2025-06-25 02:11:03 -07:00 |
|
eigen
|
20beb3702b
|
feat: add return hidden_states at async generation (#7507)
|
2025-06-25 02:10:09 -07:00 |
|
zixuanzhang226
|
f3cbd24541
|
feat: send kvmetrics from sglang scheduler (#6721)
|
2025-06-25 01:57:49 -07:00 |
|
ybyang
|
03c039c48e
|
[OAI] patch origin request_id logic (#7508)
|
2025-06-24 20:09:38 -07:00 |
|
Chang Su
|
112b496a6c
|
misc: Improvement to serving_chat.py and add more ut (#7489)
|
2025-06-24 17:19:51 -07:00 |
|
Chang Su
|
d04163b3fa
|
Fix RequestValidationError response format (#7487)
|
2025-06-23 20:35:11 -07:00 |
|
huangtingwei
|
7732bbe458
|
bugfix: Prevent global mutation of conv.stop_str across requests (#7347)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
|
2025-06-23 19:36:23 -07:00 |
|
Chang Su
|
72676cd6c0
|
feat(oai refactor): Replace openai_api with entrypoints/openai (#7351)
Co-authored-by: Jin Pan <jpan236@wisc.edu>
|
2025-06-21 13:21:06 -07:00 |
|
Keyang Ru
|
5e7fdc79fa
|
[OAI Server Refactor] [ChatCompletions & Completions] Support Return Hidden State (#7329)
Signed-off-by: keru <rukeyang@gmail.com>
|
2025-06-20 19:18:53 -07:00 |
|
Cheng Wan
|
256801e973
|
Update usage_processor.py (#7402)
|
2025-06-20 15:55:38 -07:00 |
|
yhyang201
|
dea2b84bc3
|
[OAI Server Refactor] [ChatCompletions & Completions] Implement UsageInfo Processor (#7360)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
|
2025-06-20 14:51:21 -07:00 |
|
woodx
|
cfb2fb5afc
|
[OAI refactor] Add rerank and score serving (#7399)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
|
2025-06-20 14:51:10 -07:00 |
|
Xinyuan Tong
|
0998808009
|
Refine OpenAI serving entrypoint to remove batch requests (#7372)
Signed-off-by: Xinyuan Tong <justinning0323@outlook.com>
Co-authored-by: Chang Su <csu272@usc.edu>
|
2025-06-20 14:33:43 -07:00 |
|
Ata Fatahi
|
1ab6be1b26
|
Purge VerlEngine (#7326)
Signed-off-by: Ata Fatahi <immrata@gmail.com>
|
2025-06-19 23:47:21 -07:00 |
|
woodx
|
4df5fc2156
|
Feat/refactor embedding server (#7322)
|
2025-06-19 23:46:01 -07:00 |
|
Stefan He
|
3774f07825
|
Multi-Stage Awake: Support Resume and Pause KV Cache and Weights separately (#7099)
|
2025-06-19 00:56:37 -07:00 |
|
Jinn
|
ffd1a26e09
|
Add more refactored openai test & in CI (#7284)
|
2025-06-18 13:52:55 -07:00 |
|
ishandhanani
|
31fccf5a4f
|
chore: change logs fromINFO to DEBUG for dp and add force quit for tokenizer manager (#7251)
|
2025-06-18 01:36:43 -07:00 |
|
yhyang201
|
1dffee31ac
|
OAI Server Skeleton & Core Utility Endpoints (#7179)
|
2025-06-16 20:45:55 -07:00 |
|
Xinyuan Tong
|
70c471a868
|
[Refactor] OAI Server components (#7167)
Signed-off-by: Xinyuan Tong <justinning0323@outlook.com>
|
2025-06-16 20:45:20 -07:00 |
|
woodx
|
e30ef368ab
|
Feat/support rerank (#6058)
|
2025-06-16 10:50:01 -07:00 |
|
Lianmin Zheng
|
53a525bf33
|
[Eagle] Fix kernel call after updating speculative sampling kernels (#7231)
|
2025-06-16 07:25:59 -07:00 |
|
JieXin Liang
|
ed89837cf4
|
chore: upgrade sgl-kernel v0.1.8.post2 (#7186)
Co-authored-by: zhyncs <me@zhyncs.com>
|
2025-06-14 18:26:18 -07:00 |
|
Byron Hsu
|
7d316991b2
|
[PD] Update prefill.py (#7190)
|
2025-06-14 15:59:54 -07:00 |
|
fzyzcjy
|
bec3e48402
|
Support new DeepGEMM format in per token group quant (part 2: srt) (#7155)
|
2025-06-13 14:25:40 -07:00 |
|
ishandhanani
|
f1569876d5
|
feat: add direct routing strategy to DP worker (#6884)
|
2025-06-09 11:44:05 -07:00 |
|
Yineng Zhang
|
56ccd3c22c
|
chore: upgrade flashinfer v0.2.6.post1 jit (#6958)
Co-authored-by: alcanderian <alcanderian@gmail.com>
Co-authored-by: Qiaolin Yu <qy254@cornell.edu>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
|
2025-06-09 09:22:39 -07:00 |
|
Yineng Zhang
|
23881fa60c
|
chore: upgrade sgl-kernel v0.1.6.post1 (#6957)
|
2025-06-07 17:18:55 -07:00 |
|
JieXin Liang
|
6153f2ff6e
|
chore: upgrade sgl-kernel v0.1.6 (#6945)
|
2025-06-07 02:53:26 -07:00 |
|
Chanh Nguyen
|
3f1e433903
|
Decoder-only Scoring API (#6460)
Co-authored-by: Chanh Nguyen <cnguyen@linkedin.com>
|
2025-06-04 14:14:54 -07:00 |
|
Lianmin Zheng
|
20fd53b8f6
|
Correctly abort the failed grammar requests & Improve the handling of abort (#6803)
|
2025-06-01 19:00:07 -07:00 |
|
Yineng Zhang
|
34c63731fc
|
chore: upgrade sgl-kernel v0.1.5 (#6795)
|
2025-05-31 18:32:00 -07:00 |
|
Lianmin Zheng
|
2d72fc47cf
|
Improve profiler and integrate profiler in bench_one_batch_server (#6787)
|
2025-05-31 15:53:55 -07:00 |
|
Liangsheng Yin
|
78689d3393
|
PD Rust LB (PO2) (#6437)
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
|
2025-05-29 20:50:10 +08:00 |
|
Chang Su
|
ed0c3035cd
|
feat(Tool Calling): Support required and specific function mode (#6550)
|
2025-05-23 21:00:37 -07:00 |
|
Yineng Zhang
|
0b07c4a99f
|
chore: upgrade sgl-kernel v0.1.4 (#6532)
|
2025-05-22 13:28:16 -07:00 |
|
Yineng Zhang
|
f07c6a009b
|
chore: upgrade sgl-kernel v0.1.3 (#6377)
|
2025-05-17 19:47:05 -07:00 |
|
fzyzcjy
|
f87283573e
|
Add expert distribution APIs for engine (#6290)
|
2025-05-17 18:31:51 -07:00 |
|
fzyzcjy
|
01d2838c0f
|
Fix stop_profile does not wait for finishing (#4741)
|
2025-05-17 17:06:15 -07:00 |
|
Lifu Huang
|
3cf1473a09
|
Use monotonic clock for interval measurement (#6211)
Signed-off-by: Lifu Huang <lifu.hlf@gmail.com>
|
2025-05-17 16:49:18 -07:00 |
|
Yury Sulsky
|
f19a9204cd
|
Support precomputed multimodal features for Qwen-VL and Gemma3 models. (#6136)
Co-authored-by: Yury Sulsky <ysulsky@tesla.com>
|
2025-05-16 12:26:15 -07:00 |
|
Lianmin Zheng
|
e07a6977e7
|
Minor improvements of TokenizerManager / health check (#6327)
|
2025-05-15 15:29:25 -07:00 |
|
Junrong Lin
|
f3bf611054
|
feat: add flush cache to EngineBase and HttpServerEngineAdapter (#6009)
|
2025-05-14 19:15:02 -07:00 |
|
Lianmin Zheng
|
fba8eccd7e
|
Log if cuda graph is used & extend cuda graph capture to cuda-graph-max-bs (#6201)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
|
2025-05-12 00:17:33 -07:00 |
|
fzyzcjy
|
3f2702ae51
|
Fix start_profile does not support with_stack and record_shapes (#6043)
|
2025-05-11 23:11:32 -07:00 |
|