Guy Stone
|
cd23c2f0a3
|
[Docs] add v1/score api to native api documentation (#16568)
|
2026-01-15 12:29:40 -05:00 |
|
Yi Zhong
|
d1110e1c3e
|
docs only add kimi k2 thinking and kimi linear (#15789)
|
2026-01-15 12:09:52 -05:00 |
|
shuwenn
|
9227d9f60c
|
[Docs] sort and update server_arguments.md (#17163)
|
2026-01-15 12:07:18 -05:00 |
|
Shangming Cai
|
4c59782e0f
|
Fix hybrid attention PD Disaggregation test (#17099)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-01-15 23:38:58 +08:00 |
|
cctry
|
dda35ccbd8
|
Fix gid calculation in per_tensor_absmax_kernel (#17126)
|
2026-01-15 23:21:41 +08:00 |
|
wxy
|
d11e2dc6f4
|
[diffusion] chore: improve the output_path config and enable the server to return inference duration (#16965)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2026-01-15 22:31:50 +08:00 |
|
Mick
|
16831ab6d7
|
[diffusion] fix: fix using upstream flash_attn on blackwell (#17111)
|
2026-01-15 22:30:48 +08:00 |
|
R0CKSTAR
|
c9a45b7e3c
|
[diffusion] fix: fix UMA detection (#17113)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-01-15 22:29:46 +08:00 |
|
HuangJi
|
e7df8bdc5c
|
[diffusion] refactor: move SLA to attention_backend folder (#17020)
|
2026-01-15 21:36:48 +08:00 |
|
Lancer
|
e997995037
|
[diffusion] fix: optimize text encoder CPU offload initialization to address OOM (#17064)
Signed-off-by: Lancer <maruxiang6688@gmail.com>
Co-authored-by: Lancer <maruxiang6688@gmail.com>
|
2026-01-15 21:28:57 +08:00 |
|
Hexq0210
|
6586f44ad4
|
[NPU] Add Ascend NPU best practice in doc (#17103)
|
2026-01-15 15:21:45 +08:00 |
|
Alan Kao
|
43fe3a4ddf
|
[AMD] Align alternative sgl-kernel wheel (#17092)
|
2026-01-14 21:26:45 -08:00 |
|
Bingxu Chen
|
98096b5e02
|
[AMD CI] migrate and re-enable CI tests to new CI registry (#16949)
Co-authored-by: yctseng0211 <yctseng@amd.com>
|
2026-01-14 21:25:25 -08:00 |
|
Mick
|
68e8d0f68d
|
[diffusion] CI: add testcase for cfg parallel (#17056)
|
2026-01-15 13:13:31 +08:00 |
|
sglang-bot
|
000ad42225
|
chore: bump sgl-kernel version to 0.3.21 (#17075)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-01-15 12:41:17 +08:00 |
|
Hanming Lu
|
9d5f16d456
|
[SWA] fix swa radix cache match_len_since_tombstone update when hits swa_tombstone (#17061)
|
2026-01-15 11:28:33 +08:00 |
|
Ke Bao
|
7f8a58fffb
|
Refactor prefix cache type checking (#17028)
|
2026-01-15 11:28:13 +08:00 |
|
Glen Liu
|
6b065298b5
|
[Docs] add routing-key to schedule-policy in docs (#17101)
|
2026-01-14 22:22:07 -05:00 |
|
jeff.ye
|
c020d30045
|
[Model-Gateway: grpc]: create tokenizer with chat template (#17052)
Signed-off-by: jeff.ye <jeff.ye@novita.ai>
|
2026-01-14 17:56:08 -08:00 |
|
b8zhong
|
4346db5faf
|
[Fix] Remove assertion for padding for NVFP4 weight scales to fix GLM 4.5 NVFP4 (#12497)
|
2026-01-15 08:57:14 +08:00 |
|
Артем Савкин
|
424a380077
|
[NPU] NPU quantization refactoring & more quantization formats support (#14504)
Co-authored-by: TamirBaydasov <mr.jeijy@gmail.com>
Co-authored-by: Tamir Baydasov <41994229+TamirBaydasov@users.noreply.github.com>
Co-authored-by: Савкин Артем <savkinartem@MacBook-Air-Viktoria.local>
Co-authored-by: Edward Shogulin <edward.shogulin@gmail.com>
|
2026-01-15 04:25:15 +08:00 |
|
Simo Lin
|
f091858304
|
[gRPC] Add GetLoads RPC for comprehensive load metrics (#17087)
|
2026-01-14 12:05:10 -08:00 |
|
Aurick Qiao
|
5b1215d9da
|
fix session request with None tokenizer (#16278)
|
2026-01-14 10:41:08 -08:00 |
|
Douglas Yang
|
aa2b4f7661
|
fix: renaming test file and job names + skip blocking llama4 nightly (#16971)
|
2026-01-14 09:57:59 -08:00 |
|
Simo Lin
|
b3a3f51320
|
[API] Add /v1/loads endpoint for load metrics (#16976)
|
2026-01-14 09:13:20 -08:00 |
|
shuwenn
|
de94d793ad
|
feat: support qwen3(-VL) rerank scoring&chat template (#16403)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-01-15 00:45:46 +08:00 |
|
Xiaoyu Zhang
|
0d904ef44c
|
[diffusion] fix: fix fsdp tp load make param miss parallel meta data (#17058)
|
2026-01-15 00:12:33 +08:00 |
|
Yuan Luo
|
969faaa410
|
[diffusion] fix: revise fa4 backend to support blackwell (#17077)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-01-14 23:31:46 +08:00 |
|
shuwenn
|
48c2aca9ba
|
[Env] centralize pd vars in environ.py (#16264)
|
2026-01-14 23:06:01 +08:00 |
|
fxmarty-amd
|
5af84c8af5
|
[AMD][Quantization] Add int4fp8_moe online quantization on ROCm (#7392)
Co-authored-by: Dehua Tang <dehtang@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: YC Tseng <yctseng@amd.com>
|
2026-01-14 01:44:40 -08:00 |
|
Yuan Luo
|
feae615b11
|
[VLM] Support ViT CUDA Graph for InternVL (#16732)
|
2026-01-14 17:29:23 +08:00 |
|
Netanel Haber
|
e75299a111
|
Fix issues/16714: Revert comment out of tl.debug_barrier() in causal_conv1d_triton (#16899)
Co-authored-by: Yi Zhang <1109276519@qq.com>
|
2026-01-14 17:26:48 +08:00 |
|
roikoren755
|
72bacc88c8
|
[NemotronH] Use ReplicatedLinear for fc1_latent_proj (#16569)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2026-01-14 17:08:27 +08:00 |
|
shaharmor98
|
ba625c2d90
|
Feat/support nemotron h mtp (#17013)
Signed-off-by: Shahar Mor <smor@nvidia.com>
|
2026-01-14 16:30:35 +08:00 |
|
Michael
|
b025cff441
|
[AMD] Add AMD CI registration (1-gpu unit test) to nightly CI. (#16941)
|
2026-01-13 23:41:45 -08:00 |
|
HuangJi
|
030496eb06
|
[diffusion] fix: fix --warmup-resolutions' conflict with CacheDiT (#16962)
|
2026-01-14 14:44:10 +08:00 |
|
sglang-bot
|
c86ca12875
|
chore: bump sgl-kernel version to 0.3.21 (#16888)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-01-14 13:27:49 +08:00 |
|
shuwenn
|
cd33694585
|
feat: add --admin-api-key for finer-grained endpoint auth (#15908)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
|
2026-01-13 20:21:55 -08:00 |
|
Alison Shao
|
c5e363e8e0
|
test: split Qwen3 Next tests and disable PCG tests due to intermittent failures (#16989)
|
2026-01-13 20:18:42 -08:00 |
|
Alison Shao
|
9479eca75d
|
[CI] Fix max_parallel for scheduled runs (#17046)
|
2026-01-13 20:16:32 -08:00 |
|
Mick
|
a5348eac4c
|
[diffusion] chore: avoid raising error when output resolution is not optimal (#17030)
|
2026-01-14 11:36:27 +08:00 |
|
Mick
|
9524040220
|
[diffusion] chore: refactor warmup logic (#17027)
|
2026-01-14 11:35:06 +08:00 |
|
ybyang
|
2122fea3c4
|
Update deepseekV32 Cp doc (#17054)
|
2026-01-14 11:19:26 +08:00 |
|
Liangsheng Yin
|
e2c8a50b38
|
fix grammar timeout sync across tp ranks. (#16898)
|
2026-01-14 10:26:31 +08:00 |
|
Lianmin Zheng
|
a4825ed588
|
Fix kernel type annotations for fp8 quant and logging (#16994)
|
2026-01-13 18:14:32 -08:00 |
|
Hubert Lu
|
afe285f7bd
|
[AMD] enable CUDA graph for NSA backend and fix NSA FP8 fused RMSNorm group quant (#16841)
Co-authored-by: wufann <715544327@qq.com>
|
2026-01-13 17:36:01 -08:00 |
|
Ziwen Zhao
|
cf25852a1d
|
[model-gateway] add --disable-health-check option to skip worker health probes (#17002)
|
2026-01-13 17:28:32 -08:00 |
|
Tony Lu
|
5938c3b06a
|
[model-gateway] HA - Lightweight State Layer + gRPC Mesh (#14108)
Signed-off-by: Kun(llfl) <i@imux.top>
Signed-off-by: Tony Lu <tonylu@linux.alibaba.com>
Co-authored-by: Kun(llfl) <i@imux.top>
|
2026-01-13 17:03:39 -08:00 |
|
Alison Shao
|
b880607108
|
Add 5090 dry run stage to PR test workflow (#17022)
|
2026-01-13 14:12:33 -08:00 |
|
Byron Hsu
|
339915ce2b
|
[logprob] Fix logprob + streaming for long concurrent decode by caching already processed logprob (#17005)
Co-authored-by: root <root@memx-cge-29-sr1.xpop.twttr.net>
|
2026-01-13 12:39:14 -08:00 |
|