Commit Graph

  • 5443db8759 fix: Fix AMD CI failures with HIP layernorm and PyPI connectivity (#13814) sunxxuns 2025-11-26 22:30:37 -05:00
  • 9f340ab1fb [Piecewise] support disable decode cuda graph when enable piecewise cuda graph (#13965) Stefan He 2025-11-26 18:35:59 -08:00
  • 70c6f95107 Add CODEOWNERS entry for batch_invariant_ops (#14026) Stefan He 2025-11-26 18:35:37 -08:00
  • d941a3befa Fix nightly test failure: NSA indexer dtype (#14017) alisonshao 2025-11-26 18:27:51 -08:00
  • b12c9e5c0a Fix installation for nvidia-nvshmem-cu12 (#14033) Cheng Wan 2025-11-26 18:27:12 -08:00
  • e9e90460ed [model-gateway] allow refill rate to be zero (#14030) Simo Lin 2025-11-26 17:15:27 -08:00
  • 6330d6641b Fix flashinfer cutlass MoE output shape for non-FP4-packed inputs (#14028) alisonshao 2025-11-26 17:09:02 -08:00
  • 91e8dc371a [Feat][NVFP4] Enable NVFP4 MoE for Qwen series models (eg. Qwen3-Next) #13761 (#13761) Sam 2025-11-27 08:53:45 +08:00
  • 231df4b0d4 Cleanup server args (#14027) Lianmin Zheng 2025-11-26 16:32:41 -08:00
  • b087ef8b4b Nightly test job filter (#14025) alisonshao 2025-11-26 16:04:24 -08:00
  • 5155016b56 [feat] update bucketed weights from distributed (#13824) ShawnY112358 2025-11-27 07:30:45 +08:00
  • 082b54c689 Support nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16 (and nvidia/C-RADIOv2-H) (#12277) Netanel Haber 2025-11-27 01:28:52 +02:00
  • a8ef4d1804 Add nightly test support to unified run_suite.py (#13941) alisonshao 2025-11-26 15:16:12 -08:00
  • 15ff6982b9 Support KTransformers for Qwen3-VL moe (#13983) mrhaoxx 2025-11-27 06:59:15 +08:00
  • 9adef42c60 Fix nightly test failure: CPP radix cache init (#14018) alisonshao 2025-11-26 14:58:26 -08:00
  • 685b9d82bd fix: cuda graph issue while running longcat_flash (#14007) Tianhao Zhou 2025-11-26 14:57:22 -08:00
  • a223402ffb Add adapter_model.safetensors to corruption validation for LoRA (#14022) alisonshao 2025-11-26 14:52:46 -08:00
  • 44d0a848da [model-gateway] Fix flaky test_circuit_breaker_half_open_failure_reopens (#14019) Xinyue Zhang 2025-11-26 14:52:10 -08:00
  • 5b7da0f58e Temporarily disable test_update_weights_from_disk.py in CI (#14021) alisonshao 2025-11-26 13:56:28 -08:00
  • 697a77bf73 Add stress test workflow (#13937) Douglas Yang 2025-11-26 13:24:39 -08:00
  • 0a186924ba [Code sync] Fix registration of some ops in grok & Fix oss sync scripts (#13990) Lianmin Zheng 2025-11-26 13:11:52 -08:00
  • b6312e62ea Update CODEOWNERS for layer and executor files (#14020) Stefan He 2025-11-26 12:59:04 -08:00
  • 779cbc6e4b Fix Nvidia nightly test trigger params when it is triggered by parent workflow (#13966) Kangyan-Zhou 2025-11-26 12:01:48 -08:00
  • 67c8c86722 [model-gateway][doc] Update transport terminology to protocol in README.md (#13872) Wenyi Xu 2025-11-27 03:44:59 +08:00
  • 69a03bc3f7 [ci] allow manual label to trigger ci in rust, change ci order (#14016) Simo Lin 2025-11-26 11:38:50 -08:00
  • 66f242b98f [model gateway][grpc] Add tojson filter to override minijinja's tojson (#14013) Chang Su 2025-11-26 11:05:01 -08:00
  • 8a9b8b8457 Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" (#14015) Baizhou Zhang 2025-11-26 10:45:23 -08:00
  • 7e964b5198 [ci] mark skip as success instead of failure (#14014) Simo Lin 2025-11-26 10:39:38 -08:00
  • a0d9f6cdd8 [model-gateway] fix xpu ci (#14012) Simo Lin 2025-11-26 10:33:59 -08:00
  • 8308cd3632 Support internvl on Blackwell (which doesn't support fa3): add SingletonCache support to Vision{Sdpa|Triton|Ascend}Attention (#13151) Netanel Haber 2025-11-26 20:31:45 +02:00
  • e0e8a99630 fix: correct usage of minimax-m2 deepep moe forward (#13892) Da Chen 2025-11-27 02:01:23 +08:00
  • 5102d00901 [diffusion] model: support black-forest-labs/FLUX.2-dev (#14000) Mick 2025-11-27 01:49:40 +08:00
  • 0dd759e087 Put pr-gate after check-changes (#14009) Liangsheng Yin 2025-11-27 01:40:16 +08:00
  • 5e70880e64 [model-gateway] Add PostgreSQL support to binding (#13766) Wenyi Xu 2025-11-27 01:11:13 +08:00
  • 9dab534b35 fix spec dec request level metrics (#13754) Vedant V Jhaveri 2025-11-26 09:09:21 -08:00
  • 262c3c1fde [Ascend] Support enable-mixed-chunk in non-MLA scenarios (#12491) Michelle Wu 2025-11-26 08:00:21 -08:00
  • 6c190cbda0 Rename: --hooks to --forward-hooks (#13994) Liangsheng Yin 2025-11-26 22:26:28 +08:00
  • eff6a07c8f Fix get_load API (#13991) Liangsheng Yin 2025-11-26 22:07:51 +08:00
  • b704b0a922 Optimize uneven PP layer distribution logic to improve PP performance (#13977) Shangming Cai 2025-11-26 22:05:03 +08:00
  • 15729dbc8e Use dynamically maintained num_waiting_tokens in get_load() (#13203) vipwangerxiao 2025-11-26 20:34:45 +08:00
  • 21b0582d4b [feature] Initial block diffusion language model support (#12588) Zehuan Li 2025-11-26 17:57:54 +08:00
  • 5795da5e83 [diffusion] fix: fix the issue where the qwen-edit & wan model produces incorrect output during sequence parallelism (#13922) Yuhao Yang 2025-11-26 15:58:23 +08:00
  • 540d6fee20 Support piecewise CUDA graph for embedding models (#13852) StonyPort 2025-11-26 15:29:46 +08:00
  • 5a8adca900 Turn off PREBUILD aiter in MI355 (#13963) Thomas Wang 2025-11-26 14:30:39 +08:00
  • 007c3e234c [feat] support in-flight weight update (#10071) ShawnY112358 2025-11-26 14:03:13 +08:00
  • 7130ad3a29 Fix SGLANG_ENABLE_HEALTH_ENDPOINT_GENERATION not working (#13961) fzyzcjy 2025-11-26 13:58:55 +08:00
  • 35a4c21a8a update CI permission list (#13962) Zaili Wang 2025-11-26 12:16:44 +08:00
  • 846ba3c62c Fix nightly test failures: NSA indexer dtype and CPP radix cache init (#13958) alisonshao 2025-11-25 19:52:08 -08:00
  • 18fb51583f Support FlashAttention3 page_size > 1 and topk > 1 case with paged attn and spec decode (#7725) Yubo Wang 2025-11-25 19:44:41 -08:00
  • ca5c8b16f6 [VLM] Support InternVL Vision Encoder Data Parallelism (#13925) Yuan Luo 2025-11-26 11:43:05 +08:00
  • f33e5d1ef1 fix: spec overlap predict shape does not match verify output shapes (#12786) timmy-feng 2025-11-25 19:01:06 -08:00
  • 13e5beeab4 Fix Deepseek v3.1 loading issue (#13954) Kangyan-Zhou 2025-11-25 18:28:52 -08:00
  • c53e729d45 chore: bump sgl-kernel version to 0.3.18.post1 (#13951) sglang-bot 2025-11-25 18:14:28 -08:00
  • 03a26557b7 Fix nightly-test-nvidia.yml to have the correct trigger (#13950) Kangyan-Zhou 2025-11-25 17:34:34 -08:00
  • 873382a910 [Tiny]Upgrade README for sgl-kernel (#13945) Baizhou Zhang 2025-11-25 16:46:30 -08:00
  • 391a863b3f chore: bump sgl-kernel version to 0.3.18.post1 (#13942) sglang-bot 2025-11-25 15:31:53 -08:00
  • 36b1bcd242 [chore] update torch version to 2.9 (#12969) Fan Yin 2025-11-26 06:47:34 +08:00
  • 64a11303ce Fix update weight error for blackwell DeepGEMM (#13910) fzyzcjy 2025-11-26 05:28:12 +08:00
  • 1ab6ce0e62 [Auto Sync] Improve profilers and simplify bench_one_batch_server.py (#13866) Lianmin Zheng 2025-11-25 12:13:31 -08:00
  • e99ca6ac74 Improve nightly tests (#13903) Kangyan-Zhou 2025-11-25 12:04:49 -08:00
  • fcccaf9001 Add Llama4 attention backend auto-selection (#13421) Jan Bernlöhr 2025-11-25 20:54:21 +01:00
  • 215a97fa6c Fix docstrings for v1 HiCacheStorage methods (#13851) Tova Movshovitz 2025-11-25 21:48:26 +02:00
  • 5eed5fc0b0 [DeepSeekV3.2] Centralize NSA dispatch logic in NativeSparseAttnBackend (#13544) YAMY 2025-11-25 11:32:30 -08:00
  • 808b6dfdea [Minor] Fix lint (#13938) Baizhou Zhang 2025-11-25 10:57:23 -08:00
  • 4852aa054c [misc] add llama3.1 chat template (#13935) Simo Lin 2025-11-25 09:31:54 -08:00
  • f922bfd520 [CPU] Apply PR gating rule in CI workflow (#13933) Zaili Wang 2025-11-26 00:44:16 +08:00
  • dfd7ab9682 [diffusion] feat: support LoRA (#13859) Mick 2025-11-26 00:21:33 +08:00
  • 46673b4224 [diffusion] doc: add doc for LoRA usage (#13931) Mick 2025-11-26 00:02:14 +08:00
  • 3421d049aa [CI] rename: per_commit -> registered (#13928) Liangsheng Yin 2025-11-25 23:05:06 +08:00
  • d3d404d3d7 [CI] CI registry update (#13927) Liangsheng Yin 2025-11-25 22:43:19 +08:00
  • 64225a8ae9 fix nixl prefill crash make decode health check failed (#13657) luchangli 2025-11-25 21:40:11 +08:00
  • d64bf6c6ce Support piecewise cuda graph for Qwen3-next (#13081) Chen1022 2025-11-25 21:01:27 +08:00
  • 432ecf841e [Ascend] qwen optimization (#12078) Liwansi 2025-11-25 19:44:24 +08:00
  • 0b3f002daf Update release-whl-kernel.yml (#13921) Ke Bao 2025-11-25 19:05:38 +08:00
  • 6f094deff0 [diffusion] CI: minor refactor CI for less code duplication (#13905) Mick 2025-11-25 18:44:11 +08:00
  • 59464dbf15 [Fix]: Further fix the buffer len of future map (#13916) ant-yy 2025-11-25 18:09:28 +08:00
  • 1f7fcc10d5 [diffusion] profile: fix profiling bugs (#13642) Yi Zhang 2025-11-25 17:58:35 +08:00
  • dbab5d50a3 Add test_dummy_grok_models.py to not_in_ci section (#13908) alisonshao 2025-11-25 00:51:25 -08:00
  • 760c20b360 update flashinfer_cubin==0.5.3 (#13848) Lzhang-hub 2025-11-25 16:10:34 +08:00
  • 407cb3ce1e [CI tiny fix] Enhance robustness of vision chunked prefill test with ROUGE-L metric (#13793) Xiaoyu Zhang 2025-11-25 15:41:14 +08:00
  • 7cc43bd453 Move test_dummy_grok_models.py from manual to srt (temporary) (#13901) alisonshao 2025-11-24 23:36:53 -08:00
  • cce2d748ef remove RoPE CPU fp32 tests (#13827) Zaili Wang 2025-11-25 15:22:35 +08:00
  • c1dd9a9599 [Fix] JIT kernel dependencies in other platforms (#13889) DarkSharpness 2025-11-25 15:19:17 +08:00
  • ed8786b0b9 Adding nightly tests for Kimi-K2-thinking, Qwen3, minimax-m2, GLM4.6 (#13890) Douglas Yang 2025-11-24 22:47:46 -08:00
  • f9fe06309f Fix trace publish paths in nightly-test-nvidia workflow (#13888) alisonshao 2025-11-24 21:58:26 -08:00
  • 8ff3ef1fef fix: draft model revision misuse model revision (#11893) gongwei-130 2025-11-24 21:13:37 -08:00
  • da182e4b83 [CI] fix lint error (#13891) Yibo Cai 2025-11-25 12:51:33 +08:00
  • 83e7207763 [diffusion] CI: add validation and cleanup for corrupted safetensors in multimodal loader (#13870) alisonshao 2025-11-24 20:02:43 -08:00
  • a2c388ba11 [CI] fix multimodel-gen-test job (#13874) Yibo Cai 2025-11-25 11:46:23 +08:00
  • 9384fa2729 [diffusion] refactor: remove training-related code (#13860) Mick 2025-11-25 11:38:50 +08:00
  • 173e73fa1e Fix nightly test job to fail when any test fails (#13871) alisonshao 2025-11-24 19:33:46 -08:00
  • a164259efd [Router bugfix] Fix router_manager selecting the wrong router when enable-igw. (#13572) Siyuan Chen 2025-11-25 10:51:41 +08:00
  • b0a26ba624 Add support for bf16 x bf16 cutlass fused MoE (#10275) Nicolas Castet 2025-11-24 20:49:39 -06:00
  • de430b6745 [Performance] Replace preprocess_video logic from GLM multimodal processor with transformer impl for speed up (up to 27% faster) and addressing OOM (up to 50x improvements) (#13487) Binyao Jiang 2025-11-24 18:17:13 -08:00
  • 4b45d556a7 Overlap glm moe gemms in two cuda streams (#13786) Qiaolin Yu 2025-11-24 18:15:24 -08:00
  • db0ffc09ef [NPU] Fix NPU CI (#13834) Even Zhou 2025-11-25 10:09:36 +08:00
  • eb1d885400 add LoRA warning if loading a preexisting LoRA adapter with a different name (#13822) Glen Liu 2025-11-24 18:16:41 -05:00
  • bf10869203 [Doc] Add an Introduction to Expert Parallelism (#13783) Cheng Wan 2025-11-24 14:46:51 -08:00
  • e83bd1fadc [Auto Sync] Update schedule_batch.py, schedule_policy.py, b... (20251122) (#13763) Lianmin Zheng 2025-11-24 14:33:31 -08:00
  • 9dc15d8569 fix xgrammar_backend crash with malformed inputs (#13752) gongwei-130 2025-11-24 13:54:19 -08:00