Commit Graph

  • 8a5ed2434f [NPU]support model MiniCPM3-4B for npu (#16866) McZyWu 2026-01-24 08:25:12 +08:00
  • 4c7136bb36 feature: adding openai compatible API request to bench_serving (#17219) Douglas Yang 2026-01-23 16:04:28 -08:00
  • ad05782160 fix post_residual_addition more generally (#17286) Nan Jiang 2026-01-24 07:43:37 +08:00
  • 628ab5d57b [MUSA][2/N] sgl-kernel build (#17053) R0CKSTAR 2026-01-24 06:41:47 +08:00
  • a77729a276 [MUSA][1/N] sglang.check_env (#16959) R0CKSTAR 2026-01-24 06:41:17 +08:00
  • bdaa3de075 Add return routed experts to the completions and chat/completions endpoints (#17434) Mansoor 2026-01-23 15:12:36 -05:00
  • 5438cd20ce [DLLM] Remove cuda graph batch size limitation (#17458) Tiwei Bie 2026-01-24 01:52:39 +08:00
  • 010c17a133 [Refactor] Algebraic data type for nextn config + some basic refactors (#17347) Jerry Ji 2026-01-23 09:16:55 -08:00
  • 08fcda2f63 add the fa4 mm backend and varlen func (#13539) Yi Zhong 2026-01-23 10:12:06 -05:00
  • 2fb328109f [DeepSeek V3.2] Enable trtllm NSA with bf16 kvcache (#16758) akhilg-nv 2026-01-23 04:26:21 -08:00
  • 48e9daadff Support symmetric memory pre-allocation to avoid fragmentation (#17089) Nicolas Castet 2026-01-23 03:57:04 -06:00
  • b23470e95a Fix CI install failure when rerunning tests via workflow_dispatch (#17612) Alison Shao 2026-01-23 00:04:16 -08:00
  • 17349168bb set cooldown_interval_minutes to 0 for liusy58 (#17637) siyu 2026-01-23 15:36:05 +08:00
  • 2169025b77 turn off dit_layerwise_offload for wan on rocm (#17569) Yuzhen Zhou 2026-01-22 23:22:42 -08:00
  • 69470dbc1f [NPU] update doc for Ascend NPU (#17621) Hexq0210 2026-01-23 15:21:00 +08:00
  • 69ac8b58f7 [NPU] [CI] temporarily disable mtp test (#17614) Even Zhou 2026-01-23 15:17:31 +08:00
  • 2b2f317383 fix gpt-oss launch failure with piecewise cuda graph (#17532) Minglei Zhu 2026-01-22 22:41:38 -08:00
  • d7dd0b8832 Re-enable unit-test-deepep-8-gpu and unit-test-backend-4-gpu-gb200 (#17438) Alison Shao 2026-01-22 22:31:44 -08:00
  • 56e6652d1d Lazy import torchao (#17626) Lianmin Zheng 2026-01-22 22:04:51 -08:00
  • 50a2e4345a [AMD CI] Add 2-GPU sgl-kernel Tests (#17555) Bingxu Chen 2026-01-23 13:48:52 +08:00
  • c0b5a180fe [NPU]bugfix: fix for dsv3.2 and dsvl2 (#17007) JiaruiChang5268 2026-01-23 11:15:15 +08:00
  • 7ace64d1d8 Update mamba env setting (#17566) Ke Bao 2026-01-23 11:02:32 +08:00
  • 62e6a749b0 Skip mm feature pool init to avoid EPD OOM (#16388) siyu 2026-01-23 10:53:45 +08:00
  • 04a10c9bc2 [AMD] CI - migrate perf test and fix stage-b-test-1-gpu-amd (#17340) YC Tseng 2026-01-23 10:45:05 +08:00
  • 0fec8820d1 update dependence docs of npu (#17573) amote-i 2026-01-23 09:13:36 +08:00
  • 1e8e0cca2c Update test README with CI registry documentation and 5090/H100 guidance (#17368) Alison Shao 2026-01-22 16:29:08 -08:00
  • 2399af5557 Bugfix: Writing to storage when write-back method is chosen (#14718) MMuzzammil1 2026-01-23 08:08:25 +09:00
  • 13f88045b3 configuration file support and nixl integration augmentation for hicache-storage-backend-extra-config (#16602) hxie 2026-01-22 14:31:48 -08:00
  • a296c99ff4 refactor(benchmark): prevents variable shadowing (#17607) Jacob Gordon 2026-01-22 16:00:11 -06:00
  • a921029b97 [AMD] Support ds3.2 on gfx942 platform (#17504) wufann 2026-01-23 05:57:08 +08:00
  • 15b511771d refactor(codespell): corrects typos covered up by whitelist (#17601) Jacob Gordon 2026-01-22 13:45:47 -06:00
  • 2c86a065db types: preps SessionReqNode for IDE-driven renames (#17602) Jacob Gordon 2026-01-22 13:36:23 -06:00
  • 283a2daeaa [hotfix] Reenable all reduce fusion on sm100 (#17591) Baizhou Zhang 2026-01-22 23:36:38 +08:00
  • f7a0bcda1e model: step3-vl-10b (#17513) Yuhao Yang 2026-01-22 23:15:08 +08:00
  • 72f790bf6f [BUGFIX] Skip Mamba Cache Slot 0 to Avoid Using Dummy Cache (#17404) Jincong Chen 2026-01-22 23:08:38 +08:00
  • 42523d0364 ci(pre-commit): avoids extraneous codespell exclusions (#17590) Jacob Gordon 2026-01-22 08:55:23 -06:00
  • 5324027007 [Diffusion] Make the apply_qknorm function easier to use (#17537) Xiaoyu Zhang 2026-01-22 22:32:15 +08:00
  • 3705f90629 [diffusion] model: optimize torch.compile (#17472) triple-mu 2026-01-22 23:05:47 +09:00
  • 5d299c25c0 [NPU] bugfix with Kimi-k2 and bge-reranker-v2 model (#17478) chenxu214 2026-01-22 22:02:05 +08:00
  • a4dc432587 Change naming for graph mode on multiplatform (#17469) chenxu214 2026-01-22 19:25:06 +08:00
  • 419bbcee10 refactor Qwen3-Next with a new RadixLinearAttention (#17373) Minglei Zhu 2026-01-22 01:42:06 -08:00
  • f33022d039 [RadixTree][3/N Refactor]:Support unified insert/evict params (#17401) zhangheng 2026-01-22 17:36:31 +08:00
  • 2262c5c9b5 Add ZhengWG to CI_Permission (#17572) Shangming Cai 2026-01-22 17:24:06 +08:00
  • 8dae6ec03c Add xyjixyjixyji to CI_Permission (#17559) Baizhou Zhang 2026-01-22 15:27:25 +08:00
  • 61abff66c1 [NPU] [Bug Fix] Fix typo in npu device check in gpt_oss.py (#17553) Chandrakant Khandelwal 2026-01-22 12:56:35 +05:30
  • 6a8f68b6d2 [Fix] fix device orientation for image processor (#15859) Zaili Wang 2026-01-22 15:03:10 +08:00
  • 672eb37534 [CPU][Fix CI] Solidate torch version for sgl-kernel-cpu and fix device orientation error (#17460) Zaili Wang 2026-01-22 14:04:50 +08:00
  • e2d33531f3 [Kernel] Little refactor of flashinfer allreduce norm fusion (#17474) Baizhou Zhang 2026-01-22 13:31:57 +08:00
  • 71482dd171 [diffusion] feat: enable passing Cache‑DiT config for diffusers backend (#16662) Chi McIsaac 2026-01-22 00:13:34 -05:00
  • 17807caf82 [AMD] fix amd ci dpskv32 (#17432) YC Tseng 2026-01-22 12:34:24 +08:00
  • fafa171529 [hotfix] Fixes on cuda 13 docker image (#17541) Baizhou Zhang 2026-01-22 12:29:55 +08:00
  • a5bbcda968 fix: prefer to use max_completion_tokens rather than max_tokens (#17516) Yingchun Lai 2026-01-22 12:22:20 +08:00
  • e95668abc7 [NVIDIA] Fix CUDA arch requirement in nvfp4 cast (#12581) Serge Panev 2026-01-22 05:21:11 +01:00
  • 3373545b9f [HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518) Baizhou Zhang 2026-01-22 12:17:02 +08:00
  • d6e2b88288 Add Liquid Foundation Model (LFM2) (#16890) Piotr Mazurek 2026-01-22 04:11:20 +01:00
  • f2ae066a6b Update release-branch-cut.yml for actions: write (#17539) Kangyan-Zhou 2026-01-21 18:03:27 -08:00
  • 0c2993eed0 Optimize Qwen3-VL video memory usage (#16366) cen121212 2026-01-22 09:10:08 +08:00
  • 590969ee9c [Diffusion] Support select fa2 backend in hopper (#17514) Xiaoyu Zhang 2026-01-22 08:23:53 +08:00
  • e6ccb2949b Increase wait-for-stage timeouts to handle long queue times (#17536) Alison Shao 2026-01-21 16:10:52 -08:00
  • d6ea2c529c Fix import path for UnquantizedLinearMethod in test (#17529) Alison Shao 2026-01-21 15:34:02 -08:00
  • 9be2a3a9a3 Remove test_gpt_oss_4gpu.py from __not_in_ci__ (keep in per-commit-4-gpu) (#17534) Alison Shao 2026-01-21 15:31:03 -08:00
  • b74a57a8d9 [Auto Sync] Update detokenizer_manager.py, io_struct.py, mu... (20260120) (#17442) Lianmin Zheng 2026-01-21 14:48:32 -08:00
  • 95f59c13fd [Chore] include all jit files in building packages (#17493) DarkSharpness 2026-01-22 06:48:02 +08:00
  • 85d9af51da Temporarily disable flaky test_gpt_oss_4gpu.py on B200 (#17528) Alison Shao 2026-01-21 14:01:04 -08:00
  • 858f317f13 ci(codespell): centralizes list of ignorable words (#17524) Jacob Gordon 2026-01-21 14:29:14 -06:00
  • cf89351691 [new-model] Add support for Cohere2ForCausalLM behind Command-A and Command-R Models (#16927) Lingjun Wen 2026-01-21 12:28:33 -08:00
  • 1fdf5cac39 [Auto Sync] Update environ.py, fp8.py (20260121) (#17486) Lianmin Zheng 2026-01-21 12:04:09 -08:00
  • cda43ffa4d ci: avoids duplication of codespell config (#17519) Jacob Gordon 2026-01-21 14:02:37 -06:00
  • 390898545e [Misc] Fix argument help string formatting (#17416) Yunmeng 2026-01-22 01:25:32 +08:00
  • b827e9d381 [AMD] CI - Fix sgl-kernel unittest (#17490) YC Tseng 2026-01-22 01:23:28 +08:00
  • d725487dc8 Disable swa memory for gpt-oss with spec (#17517) Ke Bao 2026-01-22 01:04:19 +08:00
  • a95c9f5b81 [NPU] Remove paged attention & Change fia to default attention (#17394) JiaruiChang5268 2026-01-21 23:58:25 +08:00
  • 19089aa431 [Diffusion] Refactor diffusion is_cuda check (#17498) Xiaoyu Zhang 2026-01-21 23:02:24 +08:00
  • 4f6f5d25c8 Support fa4 decoding (#16034) Qiaolin Yu 2026-01-21 06:54:02 -08:00
  • 458fe5a337 [docs] Show user the fastAPI docs available (#17510) Yi Zhong 2026-01-21 09:26:25 -05:00
  • 2ff0880a0e [Fix] GLM 4.7 + NVFP4 + MTP (#17166) b8zhong 2026-01-21 05:34:18 -08:00
  • 2c1b164a92 [diffusion] improve: skip negative prompt encoding when guidance_scale <= 1.0 or negative_prompt is None (#16919) Zhu Yuhua 2026-01-21 20:01:55 +08:00
  • bcc6d84f93 Use fused_sigmoid_gating_delta_rule_update_kernel for KDA (#17108) strgrb 2026-01-21 19:24:29 +08:00
  • a618202fc7 Tiny refine swa kv cache free (#17417) Ke Bao 2026-01-21 19:20:05 +08:00
  • 7520b92927 Support EPD error handling (#16670) siyu 2026-01-21 18:47:00 +08:00
  • e7224e9681 [diffusion] fix: fix the LoRA weights mismatch caused by weights packing (#17355) Fan Lin 2026-01-21 18:17:16 +08:00
  • e776239afd [diffusion] feat: support SageSparseLinearAttention attention backend (#17399) HuangJi 2026-01-21 18:13:51 +08:00
  • 1b97fa769b [BUGFIX] fix value oom in radix tree (#17400) Yi Zhang 2026-01-21 17:12:57 +08:00
  • 236772c0e1 [RadixTree][2/N Refactor]: swa cache init tiny refactor (#17397) Yi Zhang 2026-01-21 15:48:30 +08:00
  • 0d49b13fdd Fix circular import in quantization modules (#17372) Sam Shleifer 2026-01-21 02:47:09 -05:00
  • 0a7a2017a0 [diffusion] refactor: refactor and simplify teacache for cachabledit and wanvideo (#16396) blahblah 2026-01-21 15:42:45 +08:00
  • 0a9099e137 update ascend docs (#17457) amote-i 2026-01-21 15:36:26 +08:00
  • 0050c476fd Add job-level timeout for weekly test workflow (#17462) Alison Shao 2026-01-20 22:45:39 -08:00
  • 8251a74d5f [Tiny] Backward compatibility for fp4 gemm flags (#17466) Baizhou Zhang 2026-01-21 14:34:40 +08:00
  • a54d75bf2e [Fix] Set fa3 as default MHA backend on Hopper (#17425) Baizhou Zhang 2026-01-21 13:54:09 +08:00
  • 3321eb4efa Fix pr-test-finish to fail when wait-for-stage jobs fail (#17465) Alison Shao 2026-01-20 20:50:27 -08:00
  • be5121b452 Fix NSA indexer in the nightly test (#17452) Kangyan-Zhou 2026-01-20 19:58:13 -08:00
  • c3f9c30f99 [Minor] Change lora_target_modules to "all" in CI tests (#17386) Baizhou Zhang 2026-01-21 11:46:36 +08:00
  • 54a821794e [diffusion] fix: fix the bug of output_path not taking effect when generate (#17293) Fan Lin 2026-01-21 11:39:51 +08:00
  • aca354bcb3 [NPU] remove features supported on Ascend NPU (#17455) khalilzhk 2026-01-21 11:00:04 +08:00
  • aea57b33c6 [scheduler] Clear MM data of finished batch (#17251) Yinghai Lu 2026-01-20 17:45:38 -08:00
  • 823a046e8f Add hybrid parallelism test to nightly CI (#17444) Alison Shao 2026-01-20 17:43:50 -08:00
  • 648aab0ce3 Fix wait-for-stage jobs running when call-gate fails (#17443) Alison Shao 2026-01-20 17:40:05 -08:00
  • 1e309030e3 update urllib3 and gpgv Dockerfile (#17439) ishandhanani 2026-01-20 16:47:20 -06:00
  • 6092721594 [Piecewise] Fix PCG issue for multimodal and embedding model that wraps language_model (#17290) Binyao Jiang 2026-01-20 14:06:06 -08:00