Commit Graph

  • 2e14407983 Increase 5090 test parallelism from 4 to 8 (#17233) Alison Shao 2026-01-16 13:52:56 -08:00
  • d2ec128bbf fix: ci failure monitor reorganization (#17165) Douglas Yang 2026-01-16 13:25:13 -08:00
  • 7f8353aff3 [BugFix]: Fix sglang.bench_one_batch (#16925) Lingjun Wen 2026-01-16 13:15:34 -08:00
  • a7f5677abe refactor: unify registration through tokenizer_registration workflow (#17187) Chang Su 2026-01-16 10:53:36 -08:00
  • 3e968ab369 [Refactor] [CI] Remove redundant CI test runs (#17217) Makcum888e 2026-01-16 20:52:06 +03:00
  • b4fce9955a Add CI Coverage Overview workflow with detailed test listings (#16842) Alison Shao 2026-01-16 09:42:50 -08:00
  • ec9b48ea96 Add olmo3 in supported docs (#13672) Yi Zhong 2026-01-16 12:18:16 -05:00
  • a04675892e Update flashinfer to 0.6.1 (#15551) Baizhou Zhang 2026-01-17 00:48:30 +08:00
  • 82a1b645ba [DeepSeek V3.1/V3.2] Optimize fused moe configs for H20 & H20-3E based on swapab (#17133) Yongfei Xu 2026-01-17 00:10:52 +08:00
  • 6f10e17b4a [bugfix] fix qwen3-next alt_stream none issue (#17016) billishyahao 2026-01-16 22:40:25 +08:00
  • c771933dc5 [Doc] Tiny docs update for CUDA 13 (#17200) Mohammad Miadh Angkad 2026-01-16 20:53:36 +08:00
  • 9d8bbd4223 Add clear error message when OOM with symmetric memory (#17038) Nicolas Castet 2026-01-16 06:44:37 -06:00
  • daea51385d Add AFMoE model implementation (#13216) Raghav Ravishankar 2026-01-16 18:05:42 +05:30
  • 3355b6e21b feat: add request queued timeout (#17143) StonyPort 2026-01-16 17:55:09 +08:00
  • a1dd3d48ac [diffusion] hardware: support diffusion (single GPU, 3/N) (#17105) R0CKSTAR 2026-01-16 17:01:09 +08:00
  • d9ed80b9f1 fix AMD CI failure of NUMA binding (#17184) Zhiqiang Xie 2026-01-16 00:21:04 -08:00
  • 7c39ea68f3 [diffusion] model: support flux Klein (#17173) Adarsh Shirawalmath 2026-01-16 13:46:17 +05:30
  • daa4841e86 [ConfigArgumentMerger] Improve ConfigArgumentMerger compatibility with external callers (#17051) YAMY 2026-01-15 23:32:42 -08:00
  • 968c4f55b1 [AMD] Enable DeepseekV3.2 test for AMD CI (#16934) YC Tseng 2026-01-16 13:58:46 +08:00
  • 669d309a8b [model-gateway] Consolidate "unknown" model id usage (#17186) Chang Su 2026-01-15 21:02:40 -08:00
  • 21ee597e4a ci: enable offline mode when local cache is complete to avoid HF Hub … (#16121) Hudson Xing 2026-01-16 12:15:33 +08:00
  • 6ee970a365 [Diffusion] Hot fix broken output_path default value (#17180) Xiaoyu Zhang 2026-01-16 12:14:09 +08:00
  • 8ec160ed46 feature: support uvicorn access log filter(disable logging /metrics) (#15513) shuwenn 2026-01-16 12:00:06 +08:00
  • 2740ed1ae7 [eval] GSM8k support for run_eval (#17041) YAMY 2026-01-15 19:10:17 -08:00
  • d44f09ad98 [Benchmark] Add GSM8K Platinum Eval (#14565) b8zhong 2026-01-15 19:06:14 -08:00
  • e7dc85c50b Fix grammar sync across TP ranks (#17100) Lianmin Zheng 2026-01-15 18:38:01 -08:00
  • c81bad1bf7 [diffusion] feat: add cloud storage support for API (#14579) Ratish P 2026-01-16 07:59:38 +05:30
  • 0e86de7c0b Remove deepseek-r1 from THINKING_MODE_CHOICES in run_eval.py (#17178) hlu1 2026-01-15 16:53:06 -08:00
  • e3a95077bc Add dpsk-r1-fp4 in nightly perf ci (#16882) Qiaolin Yu 2026-01-15 16:13:35 -08:00
  • 146b5fcc84 [CI] Reorganize stage-b 1-GPU tests for 5090 compatibility (#16826) Alison Shao 2026-01-15 15:23:35 -08:00
  • 8b22deef5b fix【hicache】fix the KV cache resource occupation and invalid loading from prefetch when pending requests are aborted. (#16369) PiteXChen 2026-01-16 07:14:38 +08:00
  • 69822c7271 Disable unit-test-deepep-8-gpu (#17176) Alison Shao 2026-01-15 15:12:45 -08:00
  • 8b99af9af8 [Doc] Tiny update Cuda 13 environment instructions (#17174) Baizhou Zhang 2026-01-16 06:12:26 +08:00
  • 72e2f70ef7 feat(hicache): support numa detect to reduce long tail latency (#11028) JinYan Su 2026-01-16 06:11:49 +08:00
  • 3d72944fb8 [Doc] Add tip on how to use Spec V2 (#15455) b8zhong 2026-01-15 13:30:18 -08:00
  • 7dde3438e2 Show how to use cu13 image with B300 (#17170) Yi Zhong 2026-01-15 15:25:20 -05:00
  • 77fc4c4a53 Add mooncake store read/write bandwidth logs (#10598) huangtingwei 2026-01-16 04:15:41 +08:00
  • 655d2c7c2a fix: adding matrix partitioning for h200 and b200 nightly tests (#17091) Douglas Yang 2026-01-15 11:23:19 -08:00
  • 3f44268fe5 [smg] release 0.3.2 (#17168) Simo Lin 2026-01-15 11:22:18 -08:00
  • f7ec8174db [smg][ci] add make cmd to patch versions (#17167) Simo Lin 2026-01-15 11:00:53 -08:00
  • cd23c2f0a3 [Docs] add v1/score api to native api documentation (#16568) Guy Stone 2026-01-15 17:29:40 +00:00
  • d1110e1c3e docs only add kimi k2 thinking and kimi linear (#15789) Yi Zhong 2026-01-15 12:09:52 -05:00
  • 9227d9f60c [Docs] sort and update server_arguments.md (#17163) shuwenn 2026-01-16 01:07:18 +08:00
  • 4c59782e0f Fix hybrid attention PD Disaggregation test (#17099) Shangming Cai 2026-01-15 23:38:58 +08:00
  • dda35ccbd8 Fix gid calculation in per_tensor_absmax_kernel (#17126) cctry 2026-01-15 07:21:41 -08:00
  • d11e2dc6f4 [diffusion] chore: improve the output_path config and enable the server to return inference duration (#16965) wxy 2026-01-15 22:31:50 +08:00
  • 16831ab6d7 [diffusion] fix: fix using upstream flash_attn on blackwell (#17111) Mick 2026-01-15 22:30:48 +08:00
  • c9a45b7e3c [diffusion] fix: fix UMA detection (#17113) R0CKSTAR 2026-01-15 22:29:46 +08:00
  • e7df8bdc5c [diffusion] refactor: move SLA to attention_backend folder (#17020) HuangJi 2026-01-15 21:36:48 +08:00
  • e997995037 [diffusion] fix: optimize text encoder CPU offload initialization to address OOM (#17064) Lancer 2026-01-15 21:28:57 +08:00
  • 6586f44ad4 [NPU] Add Ascend NPU best practice in doc (#17103) Hexq0210 2026-01-15 15:21:45 +08:00
  • 43fe3a4ddf [AMD] Align alternative sgl-kernel wheel (#17092) Alan Kao 2026-01-15 13:26:45 +08:00
  • 98096b5e02 [AMD CI] migrate and re-enable CI tests to new CI registry (#16949) Bingxu Chen 2026-01-15 13:25:25 +08:00
  • 68e8d0f68d [diffusion] CI: add testcase for cfg parallel (#17056) Mick 2026-01-15 13:13:31 +08:00
  • 000ad42225 chore: bump sgl-kernel version to 0.3.21 (#17075) sglang-bot 2026-01-14 20:41:17 -08:00
  • 9d5f16d456 [SWA] fix swa radix cache match_len_since_tombstone update when hits swa_tombstone (#17061) Hanming Lu 2026-01-14 19:28:33 -08:00
  • 7f8a58fffb Refactor prefix cache type checking (#17028) Ke Bao 2026-01-15 11:28:13 +08:00
  • 6b065298b5 [Docs] add routing-key to schedule-policy in docs (#17101) Glen Liu 2026-01-14 22:22:07 -05:00
  • c020d30045 [Model-Gateway: grpc]: create tokenizer with chat template (#17052) jeff.ye 2026-01-15 09:56:08 +08:00
  • 4346db5faf [Fix] Remove assertion for padding for NVFP4 weight scales to fix GLM 4.5 NVFP4 (#12497) b8zhong 2026-01-14 16:57:14 -08:00
  • 424a380077 [NPU] NPU quantization refactoring & more quantization formats support (#14504) Артем Савкин 2026-01-14 23:25:15 +03:00
  • f091858304 [gRPC] Add GetLoads RPC for comprehensive load metrics (#17087) Simo Lin 2026-01-14 12:05:10 -08:00
  • 5b1215d9da fix session request with None tokenizer (#16278) Aurick Qiao 2026-01-14 10:41:08 -08:00
  • aa2b4f7661 fix: renaming test file and job names + skip blocking llama4 nightly (#16971) Douglas Yang 2026-01-14 09:57:59 -08:00
  • b3a3f51320 [API] Add /v1/loads endpoint for load metrics (#16976) Simo Lin 2026-01-14 09:13:20 -08:00
  • de94d793ad feat: support qwen3(-VL) rerank scoring&chat template (#16403) shuwenn 2026-01-15 00:45:46 +08:00
  • 0d904ef44c [diffusion] fix: fix fsdp tp load make param miss parallel meta data (#17058) Xiaoyu Zhang 2026-01-15 00:12:33 +08:00
  • 969faaa410 [diffusion] fix: revise fa4 backend to support blackwell (#17077) Yuan Luo 2026-01-14 23:31:46 +08:00
  • 48c2aca9ba [Env] centralize pd vars in environ.py (#16264) shuwenn 2026-01-14 23:06:01 +08:00
  • 5af84c8af5 [AMD][Quantization] Add int4fp8_moe online quantization on ROCm (#7392) fxmarty-amd 2026-01-14 10:44:40 +01:00
  • feae615b11 [VLM] Support ViT CUDA Graph for InternVL (#16732) Yuan Luo 2026-01-14 17:29:23 +08:00
  • e75299a111 Fix issues/16714: Revert comment out of tl.debug_barrier() in causal_conv1d_triton (#16899) Netanel Haber 2026-01-14 11:26:48 +02:00
  • 72bacc88c8 [NemotronH] Use ReplicatedLinear for fc1_latent_proj (#16569) roikoren755 2026-01-14 11:08:27 +02:00
  • ba625c2d90 Feat/support nemotron h mtp (#17013) shaharmor98 2026-01-14 10:30:35 +02:00
  • b025cff441 [AMD] Add AMD CI registration (1-gpu unit test) to nightly CI. (#16941) Michael 2026-01-13 23:41:45 -08:00
  • 030496eb06 [diffusion] fix: fix --warmup-resolutions' conflict with CacheDiT (#16962) HuangJi 2026-01-14 14:44:10 +08:00
  • c86ca12875 chore: bump sgl-kernel version to 0.3.21 (#16888) sglang-bot 2026-01-13 21:27:49 -08:00
  • cd33694585 feat: add --admin-api-key for finer-grained endpoint auth (#15908) shuwenn 2026-01-14 12:21:55 +08:00
  • c5e363e8e0 test: split Qwen3 Next tests and disable PCG tests due to intermittent failures (#16989) Alison Shao 2026-01-13 20:18:42 -08:00
  • 9479eca75d [CI] Fix max_parallel for scheduled runs (#17046) Alison Shao 2026-01-13 20:16:32 -08:00
  • a5348eac4c [diffusion] chore: avoid raising error when output resolution is not optimal (#17030) Mick 2026-01-14 11:36:27 +08:00
  • 9524040220 [diffusion] chore: refactor warmup logic (#17027) Mick 2026-01-14 11:35:06 +08:00
  • 2122fea3c4 Update deepseekV32 Cp doc (#17054) ybyang 2026-01-14 11:19:26 +08:00
  • e2c8a50b38 fix grammar timeout sync across tp ranks. (#16898) Liangsheng Yin 2026-01-14 10:26:31 +08:00
  • a4825ed588 Fix kernel type annotations for fp8 quant and logging (#16994) Lianmin Zheng 2026-01-13 18:14:32 -08:00
  • afe285f7bd [AMD] enable CUDA graph for NSA backend and fix NSA FP8 fused RMSNorm group quant (#16841) Hubert Lu 2026-01-13 17:36:01 -08:00
  • cf25852a1d [model-gateway] add --disable-health-check option to skip worker health probes (#17002) Ziwen Zhao 2026-01-13 17:28:32 -08:00
  • 5938c3b06a [model-gateway] HA - Lightweight State Layer + gRPC Mesh (#14108) Tony Lu 2026-01-14 09:03:39 +08:00
  • b880607108 Add 5090 dry run stage to PR test workflow (#17022) Alison Shao 2026-01-13 14:12:33 -08:00
  • 339915ce2b [logprob] Fix logprob + streaming for long concurrent decode by caching already processed logprob (#17005) Byron Hsu 2026-01-13 12:39:14 -08:00
  • 075c5a5789 Code clean up for fp8 quantization (#16982) Lianmin Zheng 2026-01-13 12:38:39 -08:00
  • a0b4ba9032 [diffusion] model: GLM-Image (#16894) Yuhao Yang 2026-01-14 02:02:03 +08:00
  • 2a7b67adff [CI/NPU] Fix ascend CI issue (#16953) Junrong Lin 2026-01-13 23:39:12 +08:00
  • 1d811094f8 [Misc] Auto download question file for benchmark/mtbench (#17019) elvischenv 2026-01-13 23:34:29 +08:00
  • 2ab3ed3e9e Fix sgl-kernel per_token_quant fp8 kernel scale shared_memory bug (#16886) Xiaoyu Zhang 2026-01-13 23:22:05 +08:00
  • 250477d2ac [model-gateway] Optimize L1 cache insertion with incremental hashing and tokenization (#16259) Praneth Paruchuri 2026-01-13 19:54:25 +05:30
  • af1232b2f2 [model-gateway] fix wasm example (#16924) Praneth Paruchuri 2026-01-13 19:51:47 +05:30
  • 7a869045b6 [diffusion] chore: clean excessive document (#16986) Mick 2026-01-13 21:33:50 +08:00
  • 888d7e54d1 [diffusion] fix: fix compatibility issue with torch.compile and flash attention v4 (#16790) ahb13 2026-01-13 21:11:19 +08:00
  • 3cb1fbaee4 [diffusion] fix: fix Qwen-Image-Edit Lightning LoRA alpha/rank scaling (read per-layer *.alpha) (#16935) qichu-yun 2026-01-13 21:09:11 +08:00