Commit Graph

9885 Commits

Author SHA1 Message Date
ishandhanani
498ea41ca6 dockerfile: add runtime stage + ubuntu 24.04 (#13861) 2025-12-05 00:28:36 -08:00
GMI Xiao Jin
7c744d137d [diffusion] cli: add argument --adjust-frames and --override-protected-fields (#13996)
Co-authored-by: dev <devnull@example.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-05 15:32:06 +08:00
zyksir
46b05ef58f [diffusion] fix: fix bug about pin memory when offloading (#14472) 2025-12-05 15:26:30 +08:00
Mick
beec8eed6a [diffusion] chore: further improve model searching logic (#14484) 2025-12-05 15:04:55 +08:00
Minglei Zhu
b76e303e6a clean up gemlite usage (#14444) 2025-12-04 21:52:56 -08:00
Baizhou Zhang
80a575e4e8 Add YAMY1234 to CI Permission (#14475) 2025-12-04 21:25:49 -08:00
Alison Shao
4c5074eb78 Add AMD stage support to /rerun-stage command and fix related bugs (#14463) 2025-12-04 21:13:52 -08:00
Yinghai Lu
41429a8c10 [ez] Fix typing (#14473) 2025-12-05 12:23:13 +08:00
Alison Shao
532037df78 Add CI permissions for user 'yushengsu-thu' (#14468) 2025-12-04 19:01:50 -08:00
zyksir
fa0ca97694 [diffusion] improve: further optimize model load (#13836) 2025-12-05 10:45:20 +08:00
sglang-bot
b5d3998508 Rename secrets.WHL_TOKEN -> secrets.GH_PAT_FOR_WHL_RELEASE (#14421)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2025-12-04 18:24:54 -08:00
Junrong Lin
2ecee7571c [Bug] fix not desired disable fused share experts caused by rocm logic (#14432) 2025-12-05 09:55:07 +08:00
Xinyuan Tong
6d37e70883 ministral3 (#14251)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Yueming Yuan <yy28@illinois.edu>
2025-12-04 14:31:26 -08:00
Simo Lin
c1006fd8a1 [model-gateway] grpc to leverage event type (#14450)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-12-04 14:30:19 -08:00
Simo Lin
29c6c2ea9d [model-gateway] add mistral 3 image processor (#14445)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-12-04 13:16:04 -08:00
Simo Lin
eb85fa6daf [model-gateway] move all responses api event from oai to proto (#14446)
Co-authored-by: key4ng <rukeyang@gmail.com>
2025-12-04 13:11:03 -08:00
Kay Yan
0e6441b434 fix: Correct environment variable syntax in docker-compose configuration (#8287)
Signed-off-by: Kay Yan <kay.yan@daocloud.io>
2025-12-04 13:03:39 -08:00
b8zhong
88d1bab537 add doc for quantized kv cache (#14348)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
Co-authored-by: Ho-Ren (Jack) Chuang <horenchuang@bytedance.com>
2025-12-04 13:01:05 -08:00
Sam
922756aaa1 [FIX] trtllm-moe-fp4-renorm for Qwen series models (#14350) 2025-12-04 12:52:21 -08:00
Simo Lin
d8faf2f361 [model-gateway] move oai header util to router header util (#14441)
Co-authored-by: key4ng <rukeyang@gmail.com>
2025-12-04 11:02:50 -08:00
YAMY
7dfcc78155 [DeepseekV3.2][NSA][Indexer] Fix PAGED top-k transform for NSA indexer chunked execution on H200 (#14325) 2025-12-04 10:25:03 -08:00
Simo Lin
7f3308bc68 [model-gateway] extract conversation out of oai router (#14440)
Co-authored-by: key4ng <rukeyang@gmail.com>
2025-12-04 10:16:14 -08:00
Simo Lin
fdc2ef58db [model-gateway] add llama4 vision image processor (#14438) 2025-12-04 09:52:36 -08:00
Cherry_ming
1808df48fe [NPU]add nightly-test-npu (#14143) 2025-12-05 00:43:35 +08:00
Simo Lin
b01fc161eb [model-gateway] introduce request ctx for oai router (#14434)
Co-authored-by: key4ng <rukeyang@gmail.com>
2025-12-04 08:31:44 -08:00
WenhaoZhang
788628b56f [diffusion] feat: Add Configurable Generator Device and Seed Support via API (#14366)
Co-authored-by: niehen6174 <niehen.6174@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-05 00:25:09 +08:00
Simo Lin
11d33c0e8f [model-gateway] add phi4 vision image processor (#14430) 2025-12-04 07:40:33 -08:00
Liangsheng Yin
441420e149 Add mooncake transfer_engine_bench into maunal test (#14429)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-04 22:35:00 +08:00
Raul Torres
29a2d4b59f Add 'NPU' to the runtime exception message in get_device (#14225)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2025-12-04 17:34:31 +03:00
R0CKSTAR
079ac237da [diffusion] fix: fix gen video doc (#14409)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-04 22:05:38 +08:00
Daniel Cámpora
8428078436 Add Mistral Large 3 support. (#14213)
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
Co-authored-by: Linda-Stadter <57756729+Linda-Stadter@users.noreply.github.com>
2025-12-04 20:00:05 +08:00
Xuchun Shang
af35023e65 [bug fix] fix ima with get_mla_kv_buffer_kernel overflow (#14224)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
2025-12-04 01:20:11 -08:00
Qiaolin Yu
cb8df87fc1 [1/2] Add rope kernel in sgl-kernel (#14334) 2025-12-04 16:45:44 +08:00
fzyzcjy
e3ab23c1a6 Try to fix B200 DeepEP error (#14399) 2025-12-04 00:41:24 -08:00
jianan-gu
70d2587324 [CPU] Optimize small oc GEMM for Qwen3-next on CPU (#12446)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
2025-12-04 00:38:47 -08:00
Even Zhou
894c0dc57c [NPU][1/N] NPU basic functions refactor and new modelslim quant type (#13359) 2025-12-04 16:15:31 +08:00
yctseng0211
d6c490192d [AMD] fix the regression issue for DeepseekV3 on MI300 (#14383) 2025-12-03 23:30:11 -08:00
Simo Lin
fa78c44a3d [model-gateway] introduce provider in openai router (#14394) 2025-12-03 22:47:20 -08:00
Simo Lin
654a78f9c8 [model-gateway] add phi3 vision image processor (#14381) 2025-12-03 22:32:40 -08:00
Ma Mingfei
f90b400431 [CPU] add support for mamba causal conv1d for qwen3-next (#12309) 2025-12-04 13:41:42 +08:00
Wenyi Xu
78647e08dd [model-gateway][doc] Add STDIO Explicitly to Example in README (#14393) 2025-12-03 20:36:44 -08:00
Ma Mingfei
46f21a5956 use faster covnersion from float8_e4m3fn to bfloat16 (#12316) 2025-12-04 12:34:05 +08:00
Yuan Luo
b2b09f5f24 [VLM] Introduce Cache for positional embedding ids for Qwen-VL family (#14292)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-12-04 12:32:00 +08:00
Kevin Li
04df80a9a1 Support PP x PD decode with nixl backend (#14392) 2025-12-04 12:22:17 +08:00
Zaili Wang
4f73e53dcc [CPU] document updates (#14272) 2025-12-03 19:56:06 -08:00
Qi Yuhang
16ff892c18 [sgl-kernel][Feat][B200][1/N] Support MXFP8 Grouped GEMM in Blackwell (#13731)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2025-12-04 10:09:09 +08:00
fzyzcjy
df026bb110 Fix sgl-router silently parse selector wrongly causing OME fail to discover pods (#14359) 2025-12-04 09:56:37 +08:00
Simo Lin
d42c167bfd [model-gateway] add qwen3_vl model image processor (#14377) 2025-12-03 17:16:37 -08:00
Simo Lin
388151053d [model-gateway] use worker crate in openai router (#14330) 2025-12-03 13:36:32 -08:00
b8zhong
9d82340298 Revert "Revert "enable csgmv automatically on cuda"" (#14277) 2025-12-03 13:12:30 -08:00