Commit Graph

  • 7a40e4f4a6 fix the cutlass moe tests (#10182) Rain Jiang 2025-09-08 16:24:55 -07:00
  • 19d64f2b72 fix: resolve lint issue (#10181) Yineng Zhang 2025-09-08 15:09:55 -07:00
  • a02071a12c [Bench] feat: mooncake trace integration (#9839) Teng Ma 2025-09-09 02:50:54 +08:00
  • 45b3a6a256 Revert "[ModelOpt] Fix Weight Loading for DSR1-FP4 Quantization (#9712)" (#10176) Yineng Zhang 2025-09-08 11:28:15 -07:00
  • 9a18aa54c2 [fix] Relax white space rules in EBNFComposer (#9595) LukasBluebaum 2025-09-08 19:47:19 +02:00
  • 91f0fd95a4 pref: Add H20 fp8 fused MoE kernel configs for Qwen3 (#10166) Zhiy-Zhang 2025-09-09 00:57:21 +08:00
  • 8085aca791 [Bug fix] Fix ascend mla in aclgraph (#9925) alanhe151220037 2025-09-09 00:49:43 +08:00
  • 0096798ed6 [1/2] Speed up prefill mla attention (#10156) fzyzcjy 2025-09-09 00:00:33 +08:00
  • 2c2b19b18b [CI] fix ambiguous argument in testing hybrid attentions. (#10161) Liangsheng Yin 2025-09-08 18:16:52 +08:00
  • 72f9fc5f11 Monkey patch uvicorn multi worker is_alive timeout (#10159) Liangsheng Yin 2025-09-08 17:43:23 +08:00
  • ec99668ab7 [Hicache]: Add E2E CI For 3FS-KVStore (#10131) hzh0425 2025-09-08 16:54:50 +08:00
  • 78f139812a [1/N] DP-Refactor: move communicators into tokenizer_communicator_mixin (#10028) Liangsheng Yin 2025-09-08 16:27:37 +08:00
  • bfd7a18d8d update xgrammar 0.1.24 and transformers 4.56.1 (#10155) Swipe4057 2025-09-08 11:20:31 +03:00
  • 5dd8c6444b [Bug fix] Fix Gemma 2 and fix Gemma 3 multimodal with bs > 1 on NPU (#9871) ssshinigami 2025-09-08 11:19:40 +03:00
  • ee21817c6b enable llama3.1-8B on xpu (#9434) Huaiyu, Zheng 2025-09-08 13:34:20 +08:00
  • b7d1f17b8d Revert "enable auto-round quantization model (#6226)" (#10148) Yineng Zhang 2025-09-07 22:31:11 -07:00
  • c8295d2353 enable auto-round quantization model (#6226) Weiwei 2025-09-08 13:05:35 +08:00
  • b67c277f86 [Bugfix] Qwen3MoE aclrtMemcpy failed with NPUGraph (#10013) Even Zhou 2025-09-08 12:50:49 +08:00
  • 8116804e4f Fix: (glm4v) Add missing field (#10147) Xinyuan Tong 2025-09-08 04:47:14 +00:00
  • 8c5930f08a Add speculator attention backend switch (#9981) cicirori 2025-09-08 06:44:36 +02:00
  • 3b99f23c44 [Bugfix] Retract not releasing enough memory when page size > 1 (#9989) Zhiqiang Xie 2025-09-07 21:41:50 -07:00
  • ee0b3c5bad [1/N][Bug] Fix w4afp8 MoE NaN issue (sgl-kernel, fixed) (#10108) Yuhao Yao 2025-09-08 12:39:07 +08:00
  • 6049ca209e move compile threads to an option to avoid OOM on low memory host (#10123) Rain Jiang 2025-09-07 21:36:14 -07:00
  • 7577f0e40f Add graph runner support with torch compile on CPU (#7843) Cao E 2025-09-08 12:33:58 +08:00
  • 8cda5a622c Standalone speculative decoding (#10090) Qiaolin Yu 2025-09-07 20:55:09 -07:00
  • 400d3b97ae Fix run time error in dsv3-fp8 model on mi35x (#10104) kk 2025-09-08 11:45:17 +08:00
  • 37d83c6e6d Qwen2.5-VL eagle3 infer (#8801) Lzhang-hub 2025-09-08 11:44:34 +08:00
  • 7802586cab fix the fp8 topk_config.correction_bias is none bug (#10040) Rain Jiang 2025-09-07 20:28:14 -07:00
  • bc5fc332f7 Fix slow fused add RMSNorm (#10141) fzyzcjy 2025-09-08 11:20:39 +08:00
  • f3440adcb5 vlm: enable GLM4.1V server testing & fix video processing (#10095) Xinyuan Tong 2025-09-08 02:53:08 +00:00
  • 5a7e10fe4c [MoE] fix: incorrect weight initialization for cutlass_fused_experts_fp8 (#10144) Cheng Wan 2025-09-07 19:43:59 -07:00
  • 33467c05a4 [BUG FIX] add fail check when get fail in case wait complete block (#9971) Shisong Ma 2025-09-08 09:34:04 +08:00
  • b0fcbb74d0 [DOC]: some minor updates (#10134) eigen 2025-09-07 17:58:15 -04:00
  • 76a2c86b88 Fix flashinfer version in sgl-kernel (#10135) Lianmin Zheng 2025-09-07 12:54:07 -07:00
  • e719bb0e84 [1/2] Refactor multi-tokenizer manager (#10074) Liangsheng Yin 2025-09-07 19:13:34 +08:00
  • 067246830d [Minor] fix lint in main (#10128) DarkSharpness 2025-09-07 02:36:46 -07:00
  • 617aa2b248 [Auto Sync] Update parallel_state.py (20250907) (#10126) Lianmin Zheng 2025-09-07 02:12:32 -07:00
  • 111b137964 add dataset_path for bench_one_batch_server.py (#10113) miter 2025-09-07 14:07:09 +08:00
  • 41628dc1b1 [HiCache] fix: check clear() method for storage backend (#10096) Teng Ma 2025-09-07 13:59:58 +08:00
  • a12061df4c Fix cuda graph mode in flashinfer attn backend (#10056) Ben Barsdell 2025-09-07 15:59:48 +10:00
  • 85ed8e0a5e Optimize nvfp4 block scaled gemm kernel when M is small. (#10101) Qi Yuhang 2025-09-07 13:31:00 +08:00
  • dd1e268938 CUTLASS fp8 blockwise gemm support of sm120 (#9969) Jianying 2025-09-07 13:28:54 +08:00
  • 9a7ced4e4d [Feature] LMCache Connector Integration (#9741) Yuwei An 2025-09-06 20:14:55 -07:00
  • cb3918a091 Optimize moe_sum_reduce_kernel (#9477) Yuan Luo 2025-09-07 09:16:18 +08:00
  • f3b6760213 [Auto Sync] Update server_args.py (20250906) (#10117) Lianmin Zheng 2025-09-06 16:59:36 -07:00
  • 9eb50ecc9c [router] Improve the router e2e tests (#10102) Keyang Ru 2025-09-06 16:19:28 -07:00
  • b3e7a2cee4 increase the rust e2e timeout (#10116) Keyang Ru 2025-09-06 16:17:34 -07:00
  • 00974e4f6e [CI] Refactor disaggregation tests (#10068) Shangming Cai 2025-09-06 22:14:46 +08:00
  • 5f1eb20484 [chore] Remove unused ep_moe cuda kernels (#9956) hlu1 2025-09-06 01:35:50 -07:00
  • 039cef76aa Remove non-accelerated targets(100 and up) from cmake (#10041) hlu1 2025-09-06 01:35:28 -07:00
  • 4c22ebe2e8 Disable kernel cutlass_mla_decode on SM103 (#10058) hlu1 2025-09-06 01:35:18 -07:00
  • a5a03209e9 Fix circular import (#10107) Cheng Wan 2025-09-06 01:34:17 -07:00
  • 21af5c0404 [Fix] Compatibility between DP attention and pipeline parallelism (#10100) Cheng Wan 2025-09-06 01:34:10 -07:00
  • 012584ecd5 perf: Avoid unnecessary data type conversions for DeepSeek-V3 on Blackwell (#9834) Jinyang Yuan 2025-09-06 14:06:46 +08:00
  • 90dfe3de4c [NVIDIA] disable chunked prefix cache when dp and blackwell is used (#9861) Kaixi Hou 2025-09-05 23:05:16 -07:00
  • 9a719b7afc [NVIDIA] Remove unused get_fused_moe_impl_class function (#9764) Kaixi Hou 2025-09-05 22:41:22 -07:00
  • 3fa62da78c [7/N] MoE Refactor: the implementation of new framework (#9269) Cheng Wan 2025-09-05 21:09:09 -07:00
  • dbb1235d58 [Fix] illegal sync based on undefined behaviour (#9620) DevashishLal-CB 2025-09-05 20:54:48 -07:00
  • ad26f298e2 fix double sparsity initialization (#6905) Chi-Chih Chang 2025-09-06 11:45:24 +08:00
  • 8d114f254b Fix RMSNorm API CALL mismatch issue. (#10032) sogalin 2025-09-06 11:45:13 +08:00
  • 0e78c63c0e Revert "[1/N][Bug] Fix w4afp8 MoE NaN issue (sgl-kernel) (#9953)" (#10097) Yineng Zhang 2025-09-05 19:57:53 -07:00
  • 1a3d6f31da Modify ci workflow for auto-partitioning in 2-GPU backend tests (#10029) hzh0425 2025-09-06 10:28:42 +08:00
  • 0b8c5721f1 [HiStorage] Remove delete and clear as necessary methods (#10039) Zhiqiang Xie 2025-09-05 19:27:26 -07:00
  • beac202bfd Add lora_path argument to bench_multiturn.py (#10092) Baizhou Zhang 2025-09-05 19:20:42 -07:00
  • 21b9a4b435 [router] Introduce router integration tests (#10086) Keyang Ru 2025-09-05 18:52:53 -07:00
  • db37422c92 [router] move to mcp sdk instead (#10057) Simo Lin 2025-09-05 21:03:46 -04:00
  • ab62b135c1 support Llama4 with non uniformed intermediate size across layers for… (#10047) gongwei-130 2025-09-05 17:28:15 -07:00
  • 273b28344b [Minor] Refactors KV memory pool (#9842) Xinyuan Tong 2025-09-06 00:06:08 +00:00
  • f84db115b1 Add storage read/write bandwidth logs to monitor kvcache performance (#9965) pansicheng 2025-09-06 07:52:55 +08:00
  • efb0de2c8d Update wave-lang to 3.7.0 and unify Wave kernel buffer options (#10069) jacky.cheng 2025-09-06 07:01:52 +08:00
  • 0f6ac5e21d [Bug Fix] Fix Glm4vVisionBlock norm (#9884) Adam Yanxiao Zhao 2025-09-06 05:20:36 +08:00
  • 2985090084 Update flashinfer to 0.3.1 for B300 support (#10087) hlu1 2025-09-05 13:41:01 -07:00
  • e678cc717d [bugfix]: use correct cache location for cross attention in torch native backend (#8622) Mahmoud Ashraf 2025-09-05 23:39:46 +03:00
  • 4efe844a25 enable aiter gemm_a8w8_bpreshuffle for ptpc gemm (#8555) Morpheus Guo 2025-09-06 03:54:40 +08:00
  • bde73ee43f [router] add rust cache in benchmark ci (#10080) Simo Lin 2025-09-05 12:59:36 -04:00
  • 4f0e28d7fc [router] add rust cache for rust unit test (#10079) Keyang Ru 2025-09-05 09:58:59 -07:00
  • 045ab92dc0 [router] add py binding unit tests to coverage 80% (#10043) Keyang Ru 2025-09-05 08:40:21 -07:00
  • bd7f882142 Support copying tensor from cpu to gpu without using copy engines (#10007) fzyzcjy 2025-09-05 20:07:19 +08:00
  • 5e5c30d9ab Tiny let DeepGEMM scale checks cover more cases (#7182) fzyzcjy 2025-09-05 19:52:32 +08:00
  • 9f00ec44eb Fix and enhance dumper (#8725) fzyzcjy 2025-09-05 19:51:09 +08:00
  • 8e85ee887e Support simple evals in text comparator (#8867) fzyzcjy 2025-09-05 19:50:21 +08:00
  • adf73175d6 Forbid DeepEP racing condition when too many tokens (#9567) fzyzcjy 2025-09-05 19:47:05 +08:00
  • 13705dae06 [Fix] Add speculative_draft_model_revision to server_args (#5255) DevashishLal-CB 2025-09-05 04:45:46 -07:00
  • df97b31f37 Tiny support setting numa nodes for different ranks (#10006) fzyzcjy 2025-09-05 19:01:27 +08:00
  • 339f8eef09 [1/2] Optimizations and refactors about quant kernel (#9534) fzyzcjy 2025-09-05 18:45:08 +08:00
  • afd9f2f560 Fix typo in scheduler (#9934) limingshu 2025-09-05 17:45:27 +08:00
  • f40038fb09 [Vulnerability]feat(conn): set bootstrap server host (#9931) Jimmy 2025-09-05 17:36:17 +08:00
  • bebd0576e5 Integrate trtllm ragged attention for prefill self-attention (#9801) Elfie Guo 2025-09-05 02:18:00 -07:00
  • f98366604b fix MultiTokenizerWrapper name (#10049) Huang Long 2025-09-05 13:39:46 +08:00
  • 8b3b995ac9 [router] fix release workflow to include protobuf (#10055) Chang Su 2025-09-04 22:09:30 -07:00
  • 6e95f5e5bd Simplify Router arguments passing and build it in docker image (#9964) Liangsheng Yin 2025-09-05 12:13:55 +08:00
  • 0e9387a95d fix: update gb200 dep (#10052) Yineng Zhang 2025-09-04 20:30:46 -07:00
  • fa9c82d339 chore: bump v0.5.2rc2 (#10050) Yineng Zhang 2025-09-04 20:07:27 -07:00
  • 918e3d4c27 Fix accuracy drop of dsv3 run in dp enablement (#8677) kk 2025-09-05 07:51:16 +08:00
  • e96973742c Optimized deepseek-v3/r1 model performance on mxfp4 run (#10008) kk 2025-09-05 06:11:22 +08:00
  • 93088b6975 [Hicache] Mooncake API Fix & Test, and Improved Readme (#9951) ykwd 2025-09-05 04:55:39 +08:00
  • 453511acc7 Save memory for expert model parallel (#9957) Cheng Wan 2025-09-04 13:31:47 -07:00
  • d07304870b fix 3fs zerocopy (#9938) pansicheng 2025-09-05 04:24:12 +08:00
  • b32ab0705e metrics: support customer buckets for prompt/generation_tokens_histogram (#9634) Yingchun Lai 2025-09-04 22:22:08 +08:00
  • 75ee00112d [Doc] Fix SGLang tool parser doc (#9886) Huapeng Zhou 2025-09-04 09:52:53 -04:00