almaslof
|
ff6e3ea934
|
[docs] Add missing word in argument description (#14205)
|
2025-12-06 17:56:54 -08:00 |
|
b8zhong
|
dd91d38e6a
|
[Doc] Add short explanation on page size (#14557)
|
2025-12-06 17:26:30 -08:00 |
|
Lee Nau
|
5f6f550af8
|
Update DeepSeek V3 docs to use B200 (#14447)
|
2025-12-06 17:22:11 -08:00 |
|
harrisonlimh
|
5edbe351cb
|
Update CI_PERMISSIONS.json (#14552)
|
2025-12-06 14:01:32 -08:00 |
|
sglang-bot
|
d2b42477c7
|
chore: bump sgl-kernel version to 0.3.18.post3 (#14518)
|
2025-12-06 13:15:16 -08:00 |
|
Baizhou Zhang
|
9dfa01a435
|
[Misc]Register and refactor some environs for dpsk-fp4 and DeepEp (#14538)
|
2025-12-06 12:29:16 -08:00 |
|
Hanming Lu
|
e592ee6545
|
[Qwen3-next] remove heuristics and add radix cache kl test (#14520)
|
2025-12-06 12:11:40 -08:00 |
|
Alison Shao
|
cee93a6f26
|
[Bug fix] Add /model_info endpoint to mini_lb (#14535)
|
2025-12-06 11:34:03 -08:00 |
|
Baizhou Zhang
|
bc388471d2
|
[1/n] Fix hanging during DeepGemm Warmup (#14493)
|
2025-12-06 10:44:02 -08:00 |
|
gongwei-130
|
3e40c63674
|
fix "GrammarMatcher has terminated after accepting the stop token, but is trying to find the next token mask" when both reasoning and spec are enabled (#14464)
|
2025-12-06 06:15:22 -08:00 |
|
WenhaoZhang
|
80122e4f4c
|
[diffusion] lora: fix LoRA dtype handling and weight attribute access for z-image model (#14543)
Co-authored-by: niehen6174 <nihen6174@gmail.com>
|
2025-12-06 22:14:44 +08:00 |
|
Feng Su
|
e12c6b320f
|
[model-gateway][tracing]: implement request tracing using OpenTelemetry with trace context propagation (HTTP) (#13897)
|
2025-12-06 05:59:04 -08:00 |
|
Xiaoyu Zhang
|
6d41791823
|
[diffusion] perf: add QKV fusion optimization for Flux models (#14505)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-06 20:44:16 +08:00 |
|
Mick
|
35a9a07370
|
[diffusion] refactor: simplify sampling params' override logic (#14539)
|
2025-12-06 20:23:49 +08:00 |
|
Rain Jiang
|
ea177372bd
|
support mtp with deepseek r1 nvfp4 model (#13115)
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
|
2025-12-06 00:45:54 -08:00 |
|
Baizhou Zhang
|
42fcf5438f
|
Revert "tiny remove deprecated endpoint call" (#14533)
|
2025-12-05 23:48:54 -08:00 |
|
blzheng
|
d257bf87b9
|
[CPU] add mamba fla kernels for Qwen3-next (#12324)
|
2025-12-06 14:16:23 +08:00 |
|
Yuhao Yang
|
7b0c7ad163
|
Revise DP Multi-Modal Encoder Document (#14290)
|
2025-12-06 12:56:53 +08:00 |
|
Simo Lin
|
d30d6b368f
|
[bug] fix notebook to include new keys from model_info (#14528)
|
2025-12-05 20:44:03 -08:00 |
|
Mick
|
d881f31488
|
[diffusion] chore: temporarily upgrade diffusers to make Z-image compatible with Cache-DiT (#14530)
|
2025-12-06 12:39:37 +08:00 |
|
Vincent Zhong
|
2ac5b98395
|
fix: fix rmsnorm -> layernorm in qwen3 omni (#11791)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
|
2025-12-06 12:12:57 +08:00 |
|
Hanming Lu
|
a0dde90af6
|
[CODEOWNER] update codeowner for qwen3-next related (#14522)
|
2025-12-05 16:15:10 -08:00 |
|
Alison Shao
|
b988c18eae
|
Fix safetensors validation to catch corruption after download (#14465)
|
2025-12-05 16:04:00 -08:00 |
|
Alison Shao
|
e41664ba1a
|
[Docs] Add /rerun-stage command to contribution guide (#14521)
|
2025-12-05 15:46:47 -08:00 |
|
fzyzcjy
|
3d1b591aa1
|
Tiny use trtllm_mha as default when possible (#14291)
|
2025-12-05 14:26:03 -08:00 |
|
sglang-bot
|
e11f795f63
|
chore: bump sgl-kernel version to 0.3.18.post3 (#14427)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2025-12-05 14:09:02 -08:00 |
|
Simo Lin
|
959a17466d
|
[model-gateway] fix code owner for wasm (#14516)
|
2025-12-05 13:48:25 -08:00 |
|
Simo Lin
|
b72f0268c7
|
[model-gateway] fix left over sgl-router names in wasm (#14514)
|
2025-12-05 13:16:58 -08:00 |
|
Simo Lin
|
09376fd738
|
[model-gateway] fix logs in smg workflow (#14513)
|
2025-12-05 12:50:50 -08:00 |
|
Simo Lin
|
aed835e32d
|
[model-gateway] fix left over sgl-router names to sgl-model-gateway (#14512)
|
2025-12-05 12:41:48 -08:00 |
|
Simo Lin
|
49dfa1d891
|
[model-gateway] change sgl-router to sgl-model-gateway (#14312)
|
2025-12-05 12:04:48 -08:00 |
|
Wenyi Xu
|
1ea6b740a7
|
[model-gateway] Make Tokenizer Builder Aware of Env Vars Like HF_ENDPOINT (#14405)
|
2025-12-05 10:42:06 -08:00 |
|
fzyzcjy
|
e73173b0d7
|
Fix removing worker will make it healthy forever in prometheus metrics (#14420)
|
2025-12-05 10:40:42 -08:00 |
|
Alison Shao
|
16e8463a90
|
Add Mistral Large 3 basic test to PR CI (#14460)
|
2025-12-05 10:38:52 -08:00 |
|
Simo Lin
|
cf9a774cf8
|
[model-gateway] fix server info comment (#14508)
|
2025-12-05 10:09:58 -08:00 |
|
b8zhong
|
ec7b2c16d9
|
tiny remove deprecated endpoint call (#13607)
|
2025-12-05 09:54:49 -08:00 |
|
Simo Lin
|
1569fc7f49
|
[model-gateway] reorganized conversation handler (#14507)
Co-authored-by: key4ng <rukeyang@gmail.com>
|
2025-12-05 09:31:27 -08:00 |
|
Tony Lu
|
5a46fb153d
|
[model-gateway] Add WASM support for middleware (#12471)
Signed-off-by: Tony Lu <tonylu@linux.alibaba.com>
|
2025-12-05 09:29:07 -08:00 |
|
Hudson Xing
|
38daa29466
|
Add fused FP8 KV cache write kernel for TRTLLM MHA backend (#14093)
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
|
2025-12-06 00:53:55 +08:00 |
|
blahblah
|
66984a8b3d
|
[diffusion] feat: support cache-dit integration (#14234)
Co-authored-by: shuxiguo <shuxiguo@meituan.com>
Co-authored-by: DefTruth <qiustudent_r@163.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-06 00:52:22 +08:00 |
|
roikoren755
|
889b46ea50
|
[Spec] Mamba2 support in target models (#13434)
|
2025-12-06 00:50:46 +08:00 |
|
Simo Lin
|
05284378d6
|
[model-gateway] move conversation to first class routing (#14506)
Co-authored-by: key4ng <rukeyang@gmail.com>
|
2025-12-05 08:44:13 -08:00 |
|
Mick
|
a89045603b
|
[diffusion] chore: set allowing overriding protected fields of sampling params as default behavior (#14471)
|
2025-12-06 00:22:42 +08:00 |
|
Alison Shao
|
662809874c
|
Add Mistral Large 3 to nightly CI tests (#14459)
|
2025-12-05 23:16:27 +08:00 |
|
elvischenv
|
205f041e96
|
Add Mistral Large 3 Eagle Support (#14466)
Co-authored-by: Linda-Stadter <57756729+Linda-Stadter@users.noreply.github.com>
|
2025-12-05 23:11:41 +08:00 |
|
Simo Lin
|
7235a7fbe9
|
[misc] add model arch and type to server info and use it for harmony (#14456)
|
2025-12-05 06:51:00 -08:00 |
|
Yuxuan Zhang
|
8fce9e7b2a
|
support GLM-V vision model dp (#14097)
|
2025-12-05 21:03:54 +08:00 |
|
Xiaoyu Zhang
|
5347732219
|
[diffusion] fix: Fix profiler trace missing Python stack in diffusion pipeline (#14499)
|
2025-12-05 12:12:35 +00:00 |
|
roikoren755
|
2ce121a1c3
|
Enable RadixCache for Mamba2 models (#13584)
|
2025-12-05 18:23:58 +08:00 |
|
WenhaoZhang
|
35ba6fe19e
|
[diffusion] fix: fix CLIP text encoder attention mask not used (#14364)
Co-authored-by: niehen6174 <niehen.6174@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-05 16:30:10 +08:00 |
|