yinghui
|
58095cb00a
|
Add timing metrics for requests (#12646)
Co-authored-by: Scott Lee <scottjlee@users.noreply.github.com>
|
2025-11-05 23:07:16 -08:00 |
|
gongwei-130
|
97be66c358
|
fix sgl-kernel version (#12723)
|
2025-11-05 19:01:03 -08:00 |
|
Lianmin Zheng
|
c7d57d5bb3
|
Fix CI and style (#12658)
|
2025-11-05 15:08:15 -08:00 |
|
Lianmin Zheng
|
7a21d8b276
|
Reduce the overhead of nccl symmetric memory (#12524)
Co-authored-by: Nicolas Castet <ncastet@nvidia.com>
|
2025-11-03 11:56:27 -08:00 |
|
Yineng Zhang
|
0c3543d7d5
|
chore: upgrade flashinfer 0.5.0 (#12523)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-11-02 20:54:12 -08:00 |
|
fzyzcjy
|
30ad107028
|
Try to allow NCCL cumem for multi node nvlink case (#11987)
|
2025-10-31 12:48:25 -07:00 |
|
Lianmin Zheng
|
c0652d907b
|
Clean up sgl kernel (#12413)
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
|
2025-10-31 01:13:34 -07:00 |
|
Xiaoyu Zhang
|
334543ff3b
|
Add continuous_usage_stats support for streaming responses (#12241)
|
2025-10-29 10:01:23 +08:00 |
|
Feng Su
|
ea96106000
|
[Feature] Sglang Tracing: Fine-Grained Tracking for Request Latency - Part 2 (#10804)
Signed-off-by: Feng Su <sufeng@linux.alibaba.com>
|
2025-10-28 01:25:46 -07:00 |
|
Chenxi Li
|
cc7b04a29c
|
Feature/Add GET endpoint to query loaded LoRA adapters (#12229)
|
2025-10-28 01:22:00 -07:00 |
|
satyamk7054
|
9fc3e8aac7
|
Add support for Matryoshka embeddings (#126) (#11142)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2025-10-28 02:49:36 +08:00 |
|
Chang Su
|
94aad0de99
|
[misc][grpc] Remove duplicate log (#12168)
|
2025-10-26 14:06:59 -07:00 |
|
Lianmin Zheng
|
7b36c47b3b
|
Clean up attention backend selection code & Other minor rename (#12136)
|
2025-10-25 23:50:12 -07:00 |
|
Lianmin Zheng
|
8e70064c37
|
Clean up server launch code and multi tokenizer (#12132)
|
2025-10-25 16:40:27 -07:00 |
|
Baizhou Zhang
|
4b0ac1d52a
|
Update sgl-kernel version to 0.3.16.post4 (#12125)
|
2025-10-25 14:33:33 -07:00 |
|
Teng Ma
|
96a5e4dd79
|
[Feature] Support loading weights from ckpt engine worker (#11755)
Signed-off-by: Yang Kaiyong <yangkaiyong.yky@antgroup.com>
Signed-off-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Yang Kaiyong <yangkaiyong.yky@antgroup.com>
Co-authored-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
Co-authored-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-10-23 09:23:30 -07:00 |
|
Chang Su
|
6ade6a02d4
|
[grpc] Support gRPC standard health check (#11955)
|
2025-10-22 16:59:09 -07:00 |
|
Liangsheng Yin
|
9d61205dac
|
[lint] improve ruff check (#11922)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2025-10-22 11:32:50 +08:00 |
|
Chang Su
|
70f6309cd4
|
[router][grpc] Support v1/responses API (#11926)
|
2025-10-21 17:41:48 -07:00 |
|
Yineng Zhang
|
9792b9d7e3
|
chore: upgrade flashinfer 0.4.1 (#11933)
|
2025-10-21 14:46:31 -07:00 |
|
Baizhou Zhang
|
ebff4ee648
|
Update sgl-kernel and remove fast hadamard depedency (#11844)
|
2025-10-21 13:13:54 -07:00 |
|
Zhengke Zhou
|
260fe755b6
|
Simplify multi-tokenizer (#11295)
Signed-off-by: zhengkezhou1 <madzhou1@gmail.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-10-21 16:33:29 +08:00 |
|
ybyang
|
dbb16bedd5
|
Support Thinking Budget (via custom_logit_processor for OpenAI API) [Fix #6572] (#11416)
Signed-off-by: ybyang <ybyang7@iflytek.com>
Co-authored-by: YorkSu <york_su@qq.com>
|
2025-10-21 16:27:56 +08:00 |
|
Neelabh Sinha
|
852c0578fd
|
[FEATURE] Add OpenAI-Compatible LoRA Adapter Selection (#11570)
|
2025-10-21 15:44:33 +08:00 |
|
Chang Su
|
9c0b1eb5ad
|
[router][grpc] Fix wram-up random token ids for small models (#11887)
|
2025-10-20 19:22:17 -07:00 |
|
DarkSharpness
|
276e7b3e4e
|
[Feature] New structural tag support (#10691)
|
2025-10-20 18:25:58 +08:00 |
|
ybyang
|
b5e14b2b78
|
[1/2][feature] support openai like classification api (#11618)
|
2025-10-18 19:32:48 -07:00 |
|
Lianmin Zheng
|
9eefe2c0b7
|
Set CUDA_VISIBLE_DEVICES to achieve one GPU per process (#9170)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: Cheng Wan <cwan@x.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-10-17 17:30:06 -07:00 |
|
Chang Su
|
627974405d
|
[Lint] Add python/sglang to ruff F401 checks and remove unused imports in files (#11685)
|
2025-10-17 16:49:46 -07:00 |
|
Liangsheng Yin
|
ce11dd82dc
|
[CI] Try fix broken event loop init (#11746)
|
2025-10-17 13:30:17 +08:00 |
|
Simo Lin
|
4f24ab1718
|
[router][grpc] add dissag info to warm up in grpc server (#11727)
|
2025-10-16 14:19:55 -07:00 |
|
Chang Su
|
f226d3da2a
|
Fix missing json imports in serving_responses.py (#11681)
|
2025-10-15 13:01:55 -07:00 |
|
Simo Lin
|
325951460f
|
[router][grpc] add warm up to grpc server (#11627)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
|
2025-10-14 16:11:16 -07:00 |
|
Neelabh Sinha
|
aaf7af1b17
|
[FEATURE] Add Profile Trace Merger for Distributed Traces (#11413)
|
2025-10-14 09:20:17 +08:00 |
|
Chang Su
|
887c2b4575
|
[router][grpc] Add serve_grpc to launch_server and log id for HealthCheck (#11564)
|
2025-10-13 16:07:19 -07:00 |
|
Vincent Zhong
|
a220536f40
|
[ perf ] Replace json-> orjson in hot path (#11221)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
|
2025-10-12 20:30:58 +08:00 |
|
Glen Liu
|
47c606d3dc
|
[Feature] support regex strings as a stopping condition (#10635)
|
2025-10-12 10:53:15 +08:00 |
|
Simo Lin
|
c495833186
|
[router] leverage RAII to actively cancel request during client disconnect (#11399)
|
2025-10-10 20:43:38 -04:00 |
|
Lianmin Zheng
|
9b8ebb2798
|
move more files under srt/utils (#11285)
|
2025-10-09 16:46:15 -07:00 |
|
Yineng Zhang
|
44cb060785
|
chore: upgrade flashinfer 0.4.0 (#11364)
|
2025-10-09 14:17:54 -07:00 |
|
Chang Su
|
b520958ec8
|
[router][grpc] Replace fake health check with correct ones (#11387)
|
2025-10-09 09:13:57 -07:00 |
|
Simo Lin
|
368fd20622
|
[router][grpc] disable health check generation and increase timeout (#11353)
|
2025-10-08 19:23:08 -07:00 |
|
Chang Su
|
a65ca73911
|
[router][grpc] Cleanup debug logs in grpc_server and grpc_router (#11340)
|
2025-10-08 13:26:19 -07:00 |
|
Lifu Huang
|
edefab0c64
|
[2/2] Support MHA prefill with FlashAttention 4. (#10937)
Co-authored-by: Hieu Pham <hyhieu@gmail.com>
|
2025-10-08 00:54:20 -07:00 |
|
Adarsh Shirawalmath
|
7c3f07dbcb
|
[Feature] Add /tokenize and /detokenize OpenAI compatible endpoints (#9545)
|
2025-10-08 12:38:48 +08:00 |
|
Chang Su
|
7ba3de0e92
|
[oai serving chat] Add argument --sampling-defaults and fix ChatCompletionRequest defaults (#11304)
|
2025-10-08 00:36:05 +00:00 |
|
Chang Su
|
f094e0a490
|
[router][grpc] Fix request_id extraction when n > 1 (#11311)
|
2025-10-07 19:27:56 -04:00 |
|
Chang Su
|
6f1e03a456
|
[router][grpc] Fix sampling_params.stop_strs is None (#11306)
|
2025-10-07 10:57:38 -07:00 |
|
Simo Lin
|
2fcd56eaf6
|
[router] add get server info and get model info in grpc server (#11303)
|
2025-10-07 08:36:52 -07:00 |
|
Chang Su
|
a578d300ba
|
[router][grpc] Fix proto3 default value mismatches and cleanup unused fields (#11283)
|
2025-10-06 18:54:51 -07:00 |
|