Commit Graph

9885 Commits

Author SHA1 Message Date
almaslof
ff6e3ea934 [docs] Add missing word in argument description (#14205) 2025-12-06 17:56:54 -08:00
b8zhong
dd91d38e6a [Doc] Add short explanation on page size (#14557) 2025-12-06 17:26:30 -08:00
Lee Nau
5f6f550af8 Update DeepSeek V3 docs to use B200 (#14447) 2025-12-06 17:22:11 -08:00
harrisonlimh
5edbe351cb Update CI_PERMISSIONS.json (#14552) 2025-12-06 14:01:32 -08:00
sglang-bot
d2b42477c7 chore: bump sgl-kernel version to 0.3.18.post3 (#14518) 2025-12-06 13:15:16 -08:00
Baizhou Zhang
9dfa01a435 [Misc]Register and refactor some environs for dpsk-fp4 and DeepEp (#14538) 2025-12-06 12:29:16 -08:00
Hanming Lu
e592ee6545 [Qwen3-next] remove heuristics and add radix cache kl test (#14520) 2025-12-06 12:11:40 -08:00
Alison Shao
cee93a6f26 [Bug fix] Add /model_info endpoint to mini_lb (#14535) 2025-12-06 11:34:03 -08:00
Baizhou Zhang
bc388471d2 [1/n] Fix hanging during DeepGemm Warmup (#14493) 2025-12-06 10:44:02 -08:00
gongwei-130
3e40c63674 fix "GrammarMatcher has terminated after accepting the stop token, but is trying to find the next token mask" when both reasoning and spec are enabled (#14464) 2025-12-06 06:15:22 -08:00
WenhaoZhang
80122e4f4c [diffusion] lora: fix LoRA dtype handling and weight attribute access for z-image model (#14543)
Co-authored-by: niehen6174 <nihen6174@gmail.com>
2025-12-06 22:14:44 +08:00
Feng Su
e12c6b320f [model-gateway][tracing]: implement request tracing using OpenTelemetry with trace context propagation (HTTP) (#13897) 2025-12-06 05:59:04 -08:00
Xiaoyu Zhang
6d41791823 [diffusion] perf: add QKV fusion optimization for Flux models (#14505)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-06 20:44:16 +08:00
Mick
35a9a07370 [diffusion] refactor: simplify sampling params' override logic (#14539) 2025-12-06 20:23:49 +08:00
Rain Jiang
ea177372bd support mtp with deepseek r1 nvfp4 model (#13115)
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
2025-12-06 00:45:54 -08:00
Baizhou Zhang
42fcf5438f Revert "tiny remove deprecated endpoint call" (#14533) 2025-12-05 23:48:54 -08:00
blzheng
d257bf87b9 [CPU] add mamba fla kernels for Qwen3-next (#12324) 2025-12-06 14:16:23 +08:00
Yuhao Yang
7b0c7ad163 Revise DP Multi-Modal Encoder Document (#14290) 2025-12-06 12:56:53 +08:00
Simo Lin
d30d6b368f [bug] fix notebook to include new keys from model_info (#14528) 2025-12-05 20:44:03 -08:00
Mick
d881f31488 [diffusion] chore: temporarily upgrade diffusers to make Z-image compatible with Cache-DiT (#14530) 2025-12-06 12:39:37 +08:00
Vincent Zhong
2ac5b98395 fix: fix rmsnorm -> layernorm in qwen3 omni (#11791)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
2025-12-06 12:12:57 +08:00
Hanming Lu
a0dde90af6 [CODEOWNER] update codeowner for qwen3-next related (#14522) 2025-12-05 16:15:10 -08:00
Alison Shao
b988c18eae Fix safetensors validation to catch corruption after download (#14465) 2025-12-05 16:04:00 -08:00
Alison Shao
e41664ba1a [Docs] Add /rerun-stage command to contribution guide (#14521) 2025-12-05 15:46:47 -08:00
fzyzcjy
3d1b591aa1 Tiny use trtllm_mha as default when possible (#14291) 2025-12-05 14:26:03 -08:00
sglang-bot
e11f795f63 chore: bump sgl-kernel version to 0.3.18.post3 (#14427)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2025-12-05 14:09:02 -08:00
Simo Lin
959a17466d [model-gateway] fix code owner for wasm (#14516) 2025-12-05 13:48:25 -08:00
Simo Lin
b72f0268c7 [model-gateway] fix left over sgl-router names in wasm (#14514) 2025-12-05 13:16:58 -08:00
Simo Lin
09376fd738 [model-gateway] fix logs in smg workflow (#14513) 2025-12-05 12:50:50 -08:00
Simo Lin
aed835e32d [model-gateway] fix left over sgl-router names to sgl-model-gateway (#14512) 2025-12-05 12:41:48 -08:00
Simo Lin
49dfa1d891 [model-gateway] change sgl-router to sgl-model-gateway (#14312) 2025-12-05 12:04:48 -08:00
Wenyi Xu
1ea6b740a7 [model-gateway] Make Tokenizer Builder Aware of Env Vars Like HF_ENDPOINT (#14405) 2025-12-05 10:42:06 -08:00
fzyzcjy
e73173b0d7 Fix removing worker will make it healthy forever in prometheus metrics (#14420) 2025-12-05 10:40:42 -08:00
Alison Shao
16e8463a90 Add Mistral Large 3 basic test to PR CI (#14460) 2025-12-05 10:38:52 -08:00
Simo Lin
cf9a774cf8 [model-gateway] fix server info comment (#14508) 2025-12-05 10:09:58 -08:00
b8zhong
ec7b2c16d9 tiny remove deprecated endpoint call (#13607) 2025-12-05 09:54:49 -08:00
Simo Lin
1569fc7f49 [model-gateway] reorganized conversation handler (#14507)
Co-authored-by: key4ng <rukeyang@gmail.com>
2025-12-05 09:31:27 -08:00
Tony Lu
5a46fb153d [model-gateway] Add WASM support for middleware (#12471)
Signed-off-by: Tony Lu <tonylu@linux.alibaba.com>
2025-12-05 09:29:07 -08:00
Hudson Xing
38daa29466 Add fused FP8 KV cache write kernel for TRTLLM MHA backend (#14093)
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
2025-12-06 00:53:55 +08:00
blahblah
66984a8b3d [diffusion] feat: support cache-dit integration (#14234)
Co-authored-by: shuxiguo <shuxiguo@meituan.com>
Co-authored-by: DefTruth <qiustudent_r@163.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-06 00:52:22 +08:00
roikoren755
889b46ea50 [Spec] Mamba2 support in target models (#13434) 2025-12-06 00:50:46 +08:00
Simo Lin
05284378d6 [model-gateway] move conversation to first class routing (#14506)
Co-authored-by: key4ng <rukeyang@gmail.com>
2025-12-05 08:44:13 -08:00
Mick
a89045603b [diffusion] chore: set allowing overriding protected fields of sampling params as default behavior (#14471) 2025-12-06 00:22:42 +08:00
Alison Shao
662809874c Add Mistral Large 3 to nightly CI tests (#14459) 2025-12-05 23:16:27 +08:00
elvischenv
205f041e96 Add Mistral Large 3 Eagle Support (#14466)
Co-authored-by: Linda-Stadter <57756729+Linda-Stadter@users.noreply.github.com>
2025-12-05 23:11:41 +08:00
Simo Lin
7235a7fbe9 [misc] add model arch and type to server info and use it for harmony (#14456) 2025-12-05 06:51:00 -08:00
Yuxuan Zhang
8fce9e7b2a support GLM-V vision model dp (#14097) 2025-12-05 21:03:54 +08:00
Xiaoyu Zhang
5347732219 [diffusion] fix: Fix profiler trace missing Python stack in diffusion pipeline (#14499) 2025-12-05 12:12:35 +00:00
roikoren755
2ce121a1c3 Enable RadixCache for Mamba2 models (#13584) 2025-12-05 18:23:58 +08:00
WenhaoZhang
35ba6fe19e [diffusion] fix: fix CLIP text encoder attention mask not used (#14364)
Co-authored-by: niehen6174 <niehen.6174@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-05 16:30:10 +08:00