Kangyan-Zhou
|
592603d77b
|
Fix flaky streaming logprobs test by handling detokenizer text buffering (#17687)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-01-25 15:09:06 -08:00 |
|
Mansoor
|
bdaa3de075
|
Add return routed experts to the completions and chat/completions endpoints (#17434)
|
2026-01-23 12:12:36 -08:00 |
|
Yingchun Lai
|
a5bbcda968
|
fix: prefer to use max_completion_tokens rather than max_tokens (#17516)
|
2026-01-21 20:22:20 -08:00 |
|
Qiaolin Yu
|
76b06bee03
|
[New Model] GLM4.7-Flash (#17247)
Co-authored-by: zRzRzRzRzRzRzR <2448370773@qq.com>
Co-authored-by: JustinTong0323 <justinning0323@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-01-20 23:44:16 +08:00 |
|
ybyang
|
ebca5879a1
|
Fix v32 continue_final_message not work (#16567)
|
2026-01-19 21:45:39 +08:00 |
|
Lianmin Zheng
|
e7dc85c50b
|
Fix grammar sync across TP ranks (#17100)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2026-01-15 18:38:01 -08:00 |
|
shuwenn
|
de94d793ad
|
feat: support qwen3(-VL) rerank scoring&chat template (#16403)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-01-15 00:45:46 +08:00 |
|
fzyzcjy
|
15da3061bc
|
Tiny pass routing key to scheduler processes (#16839)
|
2026-01-10 10:13:00 +08:00 |
|
yuhao
|
d2ea44f775
|
VLM: enhance VL embedding model with video input support and revise warm-up strategy (#16635)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-01-08 22:12:01 +08:00 |
|
Harish
|
156d97b219
|
Fix KeyError when logprobs=false in completions endpoint (#16095)
|
2026-01-07 15:49:02 -08:00 |
|
Yuan Luo
|
5f3eb377e0
|
[VLM] Support request level max_dynamic_patch for OpenAI request (#16268)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-01-04 13:04:43 +08:00 |
|
Kangyan-Zhou
|
a91e072f33
|
Use allow auto truncate in the OpenAI API endpoint (#15369)
|
2025-12-25 18:45:57 -08:00 |
|
Qiaolin Yu
|
e220da1723
|
tiny fix sampling seed for completion api (#15498)
|
2025-12-19 20:32:13 -08:00 |
|
Jimmy
|
216067c0cb
|
feat(dsv32): better error handling for DeepSeek-v3.2 encoder (#14353)
|
2025-12-18 16:34:27 -08:00 |
|
Lianmin Zheng
|
d1f0063262
|
Clean up __init__ function of the scheduler and event loop for PD (#15298)
|
2025-12-18 01:35:14 -08:00 |
|
Kenny Yao
|
5e96beb3e5
|
Adding tool calling and reasoning parser support for Intern-S1 (#14866)
Co-authored-by: KennyYao2001 <kennyy@andrew.cmu.edu>
Co-authored-by: Cursor AI <cursor@example.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Changyi Yang <112288487+ChangyiYang@users.noreply.github.com>
|
2025-12-16 00:01:58 -08:00 |
|
Muqi Li
|
1da5cd638f
|
[Bugfix][Tool Call] Add null system prompt to support tool system prompt (#15092)
Signed-off-by: Muqi Li <muqi1029@gmail.com>
|
2025-12-15 19:40:10 -08:00 |
|
danielafrimi
|
8102e36b5d
|
Add NanoV3 reasoning parser support (#15113)
|
2025-12-14 12:09:44 -08:00 |
|
Xinyuan Tong
|
96705514bd
|
fix: dpskv32 chat history processing, default drop_thinking to true (#15064)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-13 12:16:04 -08:00 |
|
Xinyuan Tong
|
8bf10e71cd
|
remove dpsk3.2 sys prompt (#14923)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-11 20:26:13 -08:00 |
|
Xinyuan Tong
|
4885f8b9c1
|
fix: handle Jinja2 template errors as client errors in OpenAIServingChat (#14748)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-11 17:06:26 -08:00 |
|
Xinyuan Tong
|
9975acf50f
|
[refactor] Update reasoning parameter to require_reasoning (#14922)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-11 16:55:59 -08:00 |
|
Jimmy
|
5c961756a3
|
[FIX][DS32]openai protocol: support openai message role: developer (#14304)
|
2025-12-11 11:34:41 -08:00 |
|
Liangsheng Yin
|
cbc7dcdaa7
|
Re-add the API serving timing metrics. (#14744)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
Co-authored-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
|
2025-12-10 10:17:48 +09:00 |
|
Lianmin Zheng
|
ab0048793c
|
Revert "[Feat] Add received_time in serving_base" (#14743)
|
2025-12-09 06:52:27 -08:00 |
|
zhanghaotong
|
fe7f91ef82
|
[Feat] Add received_time in serving_base (#13432)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
|
2025-12-09 22:00:24 +09:00 |
|
fzyzcjy
|
da3dc497b0
|
Fix dp-aware incompatible with completions and chat completions APIs (#14647)
|
2025-12-09 13:24:53 +08:00 |
|
Muqi Li
|
06836ad02a
|
[Reasoning + Structured Output] make reasoning compatible with structured output (#12551)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-08 01:28:39 -08:00 |
|
Eva20150932-atlascloud
|
7c38eca1e4
|
feat: DeepSeek new v3.2 encoding (#14249)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-02 11:41:05 -08:00 |
|
Lianmin Zheng
|
155a9e7237
|
Fix condition for streaming output_ids in tokenizer manager (#13759)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Chang Su <chang.s.su\n@oracle.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-11-29 13:56:15 -08:00 |
|
Binyao Jiang
|
26ca07469b
|
[GLM4.6v] Relax the constraint of non-user role chat completion message schema for new GLM-v release (#13258)
|
2025-11-18 10:58:17 +08:00 |
|
Binyao Jiang
|
90c18a16cb
|
[GLM4.6v] Required changes for bumping up to transformer 5.x (#13229)
|
2025-11-18 10:58:00 +08:00 |
|
Frank Minors
|
7f5055ed5e
|
Fix cached tokens usage bug (#12814)
Co-authored-by: FrankMinions <liuchen@shinemo.com>
|
2025-11-12 01:02:43 +08:00 |
|
yinghui
|
58095cb00a
|
Add timing metrics for requests (#12646)
Co-authored-by: Scott Lee <scottjlee@users.noreply.github.com>
|
2025-11-05 23:07:16 -08:00 |
|
Lianmin Zheng
|
c0652d907b
|
Clean up sgl kernel (#12413)
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
|
2025-10-31 01:13:34 -07:00 |
|
Xiaoyu Zhang
|
334543ff3b
|
Add continuous_usage_stats support for streaming responses (#12241)
|
2025-10-29 10:01:23 +08:00 |
|
Chenxi Li
|
cc7b04a29c
|
Feature/Add GET endpoint to query loaded LoRA adapters (#12229)
|
2025-10-28 01:22:00 -07:00 |
|
satyamk7054
|
9fc3e8aac7
|
Add support for Matryoshka embeddings (#126) (#11142)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2025-10-28 02:49:36 +08:00 |
|
Liangsheng Yin
|
9d61205dac
|
[lint] improve ruff check (#11922)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2025-10-22 11:32:50 +08:00 |
|
Chang Su
|
70f6309cd4
|
[router][grpc] Support v1/responses API (#11926)
|
2025-10-21 17:41:48 -07:00 |
|
ybyang
|
dbb16bedd5
|
Support Thinking Budget (via custom_logit_processor for OpenAI API) [Fix #6572] (#11416)
Signed-off-by: ybyang <ybyang7@iflytek.com>
Co-authored-by: YorkSu <york_su@qq.com>
|
2025-10-21 16:27:56 +08:00 |
|
Neelabh Sinha
|
852c0578fd
|
[FEATURE] Add OpenAI-Compatible LoRA Adapter Selection (#11570)
|
2025-10-21 15:44:33 +08:00 |
|
DarkSharpness
|
276e7b3e4e
|
[Feature] New structural tag support (#10691)
|
2025-10-20 18:25:58 +08:00 |
|
ybyang
|
b5e14b2b78
|
[1/2][feature] support openai like classification api (#11618)
|
2025-10-18 19:32:48 -07:00 |
|
Chang Su
|
f226d3da2a
|
Fix missing json imports in serving_responses.py (#11681)
|
2025-10-15 13:01:55 -07:00 |
|
Vincent Zhong
|
a220536f40
|
[ perf ] Replace json-> orjson in hot path (#11221)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
|
2025-10-12 20:30:58 +08:00 |
|
Glen Liu
|
47c606d3dc
|
[Feature] support regex strings as a stopping condition (#10635)
|
2025-10-12 10:53:15 +08:00 |
|
Adarsh Shirawalmath
|
7c3f07dbcb
|
[Feature] Add /tokenize and /detokenize OpenAI compatible endpoints (#9545)
|
2025-10-08 12:38:48 +08:00 |
|
Chang Su
|
7ba3de0e92
|
[oai serving chat] Add argument --sampling-defaults and fix ChatCompletionRequest defaults (#11304)
|
2025-10-08 00:36:05 +00:00 |
|
Vincent Zhong
|
36a6b8dbfc
|
Update v1/responses to be more OpenAI-compatible. (#9624)
|
2025-10-05 18:47:46 +00:00 |
|