Commit Graph

  • 94bcc19bce [VLM] Support Video for InternVL3_5 (#15942) Yuan Luo 2025-12-30 17:07:49 +08:00
  • db3821a9ef [PP] Add a minimum chunk value for PP dynamic chunking (#16140) Shangming Cai 2025-12-30 16:35:39 +08:00
  • 8c5d91b884 Fix race condition in /tag-and-rerun-ci command (#16142) Alison Shao 2025-12-29 23:55:49 -08:00
  • 6a3e709253 Tiny rename test_deepseek_v3_fp4_mtp_stage_b.py (#16141) Qiaolin Yu 2025-12-29 23:55:07 -08:00
  • c2601f0d21 Fix wrong assigning extend_input_len_per_req with eagle. (#16129) Liangsheng Yin 2025-12-30 15:51:53 +08:00
  • 1048803c1f Reworked fast_pos_embed_interpolate() using torch (#10959) Vitaly Tuzov 2025-12-30 09:45:34 +03:00
  • 8a84b1e7e0 [diffusion] model: support TurboWan2.1-T2V-1.3B/14B SLA (#15888) HuangJi 2025-12-30 14:30:22 +08:00
  • 5e20e7a60d [AMD CI] Organize AMD nightly perf test files (#16114) Bingxu Chen 2025-12-30 14:09:46 +08:00
  • 3946dad6fa Add attack204 into CI_PERMISSION (#16131) Liangsheng Yin 2025-12-30 13:55:00 +08:00
  • ac03ec0851 Add PR review process into template. (#16133) Liangsheng Yin 2025-12-30 13:45:56 +08:00
  • 9416464682 Fix Qwen Next GDN w/ Radix Cache (#16053) Stefan He 2025-12-29 21:06:06 -08:00
  • 5fb734f1a5 Enhance comments in set_extend_input_len method (#16130) Cheng Wan 2025-12-29 20:47:33 -08:00
  • 60f1ca6925 Refactor: Moving extend_logprob_start_len calculation out of prepare_for_extend (#16105) Cheng Wan 2025-12-29 20:38:33 -08:00
  • 9e263c2162 [LoRA] Torch native backend: rework implementation and updated tests (#15187) Vladimir Serov 2025-12-30 06:48:51 +03:00
  • 0d003e34b0 Reduce CI failure monitor to run once every 12 hours (#16123) Kangyan-Zhou 2025-12-29 19:44:25 -08:00
  • f253f43c9d [diffusion] CI: fix LoRA downloading issues and respect offline flag (#15813) Prozac614 2025-12-30 11:39:27 +08:00
  • 1e45320198 [diffusion] improve: tiny improve layerwise offload manager by consolidating weights per layer (#16081) Mick 2025-12-30 11:31:00 +08:00
  • 26e17f9076 [diffusion] improve: tiny speedup qwen-image-edit-2511 by avoiding unnecessary calculation (#15896) Mick 2025-12-30 10:10:32 +08:00
  • 3de23274ee Clean up swa handling in fa3 backend (#15877) Ke Bao 2025-12-30 08:51:40 +08:00
  • f4ec6f8e17 ci: migrate remaining spec/eagle tests to test/registered/spec/ (#15800) Alison Shao 2025-12-29 14:00:00 -08:00
  • 9c4eb46099 Add a new branch cut GH workflow, and adopt setuptools-scm for version control (#15985) Kangyan-Zhou 2025-12-29 13:51:21 -08:00
  • c2e0913e17 Fix extend_input_len calculation in decode.py (#16103) Cheng Wan 2025-12-29 13:13:24 -08:00
  • 8c6f865afb [model-gateway]: optimize prefix_match with zero-copy tenant and deferred char count (#16099) Simo Lin 2025-12-29 11:32:57 -08:00
  • 269aa27bff Enable testing slash command handler changes on non-fork PRs (#15921) Alison Shao 2025-12-29 11:30:39 -08:00
  • f39382c6ca [model-gateway] Fix duplicate classify prefix in response ID (#16101) Simo Lin 2025-12-29 11:07:21 -08:00
  • 2ff289e2dc [model-gateway] Generate UUID-based request IDs for embedding/classify (#16100) Simo Lin 2025-12-29 10:49:58 -08:00
  • b2a3f0551a [model-gateway] Wire classify pipeline to gRPC router (#16098) Simo Lin 2025-12-29 10:44:45 -08:00
  • 684e148e38 [model-gateway] Optimize INSERT with leaf-only timestamp updates (#16097) Simo Lin 2025-12-29 10:43:20 -08:00
  • f44c4b3715 [model-gateway] Add classify pipeline stages and protocol types (#16094) Simo Lin 2025-12-29 10:23:36 -08:00
  • 7380ec9d55 [CI] fix test_mla_deepseek_v3.py (#16096) shuwenn 2025-12-30 02:18:02 +08:00
  • ac78f96ec5 [model-gateway] Optimize radix tree timestamp updates for multi-tenant scaling (#16093) Simo Lin 2025-12-29 10:08:19 -08:00
  • 278012caa0 [Feature] support bench jsonl files with sharegpt format (#15057) jiapingW 2025-12-30 02:06:11 +08:00
  • 7f5879985c [model-gateway] Improve tree benchmark with realistic multi-tenant scenarios (#14838) Simo Lin 2025-12-29 08:37:24 -08:00
  • 162d1cf9be [model-gateway] Add classification model support infrastructure (#16061) Simo Lin 2025-12-29 08:34:05 -08:00
  • 8e08207c18 [diffusion] fix: fix serving with dit-layerwise-offload enabled (#16066) Mick 2025-12-30 00:21:42 +08:00
  • c31f62722c [model-gateway] fix tokenizer to match transformers special token handling (#16087) Simo Lin 2025-12-29 08:13:03 -08:00
  • b5d9fc873b [diffusion] chore: minor refactor by streamlining the VAE class hierarchy (#16069) Mick 2025-12-29 23:37:59 +08:00
  • 88f3de2514 [JIT kernel] Jit kernel add codeowners (#16085) Xiaoyu Zhang 2025-12-29 22:32:55 +08:00
  • 24616c5234 [Diffusion] Qwen image edit support qknorm optimization (#16062) Xiaoyu Zhang 2025-12-29 21:37:24 +08:00
  • d48723b77d Clamp logprob tokens with model vocab size (#14414) cklxx 2025-12-29 20:37:07 +08:00
  • f3d73b0199 [Diffusion] Refactor qwen_image's rope in a single helper func (#16047) Xiaoyu Zhang 2025-12-29 17:24:26 +08:00
  • 4ab66d956f [HiCache] Fix deadlock when creating new group (#15805) Xuchun Shang 2025-12-29 17:13:04 +08:00
  • de2799f3f5 [docs] Fix non-clickable ToC links in model gateway documentation (#16054) Simo Lin 2025-12-28 23:35:37 -08:00
  • f784cbfa92 Update model and feature support for Ascend NPU (#16003) Hexq0210 2025-12-29 15:16:25 +08:00
  • 097330909f [model-gateway] update WorkerRegistryStats with connection mode and circuit breaker info (#16046) Simo Lin 2025-12-28 23:14:57 -08:00
  • a44d0079f8 [ci] update genai bench to 0.0.3 for pd testing (#16051) Simo Lin 2025-12-28 23:14:44 -08:00
  • 8305dc1718 [Diffusion] Disable packed QKV for FLUX & Z-Image (#16038) Xiaoyu Zhang 2025-12-29 14:33:05 +08:00
  • ec8c831d5a [model-gateway]: optimize metrics for minimal CPU and memory overhead (#16041) Simo Lin 2025-12-28 22:28:08 -08:00
  • f13949e5a9 [model-gateway] perf: optimize observability logging for minimal CPU/memory overhead (#16039) Simo Lin 2025-12-28 21:20:46 -08:00
  • c236a3fdf0 [model-gateway][CI] Display benchmark results in GitHub Actions summary (#16037) Simo Lin 2025-12-28 20:57:27 -08:00
  • 41a1d16b02 [model-gateway] Organize CLI arguments into logical groups for better --help output (#16035) Simo Lin 2025-12-28 20:25:53 -08:00
  • 9884c9fd4c [model-gateway] Organize Rust CLI arguments into logical groups for better --help output (#16036) Simo Lin 2025-12-28 20:25:38 -08:00
  • a435f55d18 Tiny print launch command with shlex (#16010) Liangsheng Yin 2025-12-29 11:26:46 +08:00
  • 2ec6fa3c54 feat PD: add eagle3 support for DeepSeek V3 in EP mode (#14280) Mike Qiu 2025-12-29 11:12:34 +08:00
  • c58a573a40 [Docs] Improve documentation index page (#16028) Lianmin Zheng 2025-12-28 18:44:52 -08:00
  • b840d6aaeb [diffusion] webui: reference to content task and better visualization capabilities (#16017) Li Jinliang 2025-12-29 10:22:21 +08:00
  • ef4b3c0e96 Add host tensor allocator for memory_pool_host and support Mooncake standalone storage (#14873) EkiRui 2025-12-29 08:28:50 +08:00
  • 6f9d0a89a0 [scheduler] fix: correcting extend_logprob_start_len calculation (#15922) Cheng Wan 2025-12-28 14:57:04 -08:00
  • d7a3336ebe [diffusion] fix: fix stages not logged when perf_dump_path is provided (#16016) Mick 2025-12-28 23:17:43 +08:00
  • 7d02c8e59f [JIT kernel] CI support jit kernel tests (#15939) Xiaoyu Zhang 2025-12-28 23:11:02 +08:00
  • e6d5a213ad Fix metrics (#15998) Lianmin Zheng 2025-12-28 05:03:49 -08:00
  • 9f8e23071a [diffusion] chore: fix default offload setting for image generation model (#15928) Mick 2025-12-28 20:45:33 +08:00
  • d90f9bfc4e Temporarily disable temp_prefill_info assertion to unblock CI (#16008) fzyzcjy 2025-12-28 20:13:06 +08:00
  • 3881bc8d0b [diffusion] CI: relax threshold by supporting different profiles (#16002) Mick 2025-12-28 20:05:07 +08:00
  • 208e6a9dac [Doc]Update MTP moe backends for EP document (#16013) Baizhou Zhang 2025-12-28 19:36:23 +08:00
  • be3828a13b Tiny cleanup duplicate code for multi-layer eagle worker. (#16004) Liangsheng Yin 2025-12-28 18:08:20 +08:00
  • 5969be2f06 Apply fixture-kit mode to MMMUVLMMixin (#15615) lif 2025-12-28 17:22:05 +08:00
  • 8fab48952e Tiny fix WASM test errors on machines with many cores (#15992) fzyzcjy 2025-12-28 17:12:04 +08:00
  • 7ccaec6467 Add micro benchmarks for manual policy (#15991) fzyzcjy 2025-12-28 17:11:56 +08:00
  • c457aad54a Update test parameters for deepep_large test (#16001) Cheng Wan 2025-12-28 00:58:19 -08:00
  • c7e7bfa32c Add EAGLE3 test with MMLU dataset. (#15945) Liangsheng Yin 2025-12-28 16:07:58 +08:00
  • bf90ea9c5b Unify spec v2's naming manner. (#15990) Liangsheng Yin 2025-12-28 14:14:52 +08:00
  • 26c5091217 Tiny extract PeriodicTask in router (#15988) fzyzcjy 2025-12-28 14:07:13 +08:00
  • 325a4c1945 Tiny add smg_manual_policy_cache_entries metric (#15987) fzyzcjy 2025-12-28 14:07:07 +08:00
  • 0294844f04 [fix]deepgemm precompile when warmup (#15891) TZHelloWorld 2025-12-28 13:47:39 +08:00
  • 0e536600e8 Refactor: separate CI-specific weight validation into dedicated module (#15216) Alison Shao 2025-12-27 20:50:39 -08:00
  • d70c265533 SGLang Tracing: fix attribute errors (header extraction & bootstrap span closing) (#15693) Vladislav Nosivskoy 2025-12-28 07:44:31 +03:00
  • 656f4d69a1 Refactor fp8 nextn layer for DeepSeek nvfp4 checkpoint (#15353) Baizhou Zhang 2025-12-28 11:57:09 +08:00
  • 8e43980ebb [Feature] JIT Fused QK norm + qk norm clean up (#15835) DarkSharpness 2025-12-28 11:53:50 +08:00
  • 474a4699c5 Add correctness validation for decode_attention test (#15806) Hudson Xing 2025-12-28 11:46:17 +08:00
  • b4a00ed2d9 [diffusion] chore: clean ComposedPipelineBase (#15937) Mick 2025-12-28 11:43:25 +08:00
  • f55d608c89 Tiny fix cannot launch nvfp4 checkpoint with bf16 kv cache (#15986) fzyzcjy 2025-12-28 11:22:20 +08:00
  • 349ce2dd19 Support kv8 (FP8) with torch_native attention backend (#12596) Ho-Ren (Jack) Chuang 2025-12-27 18:48:30 -08:00
  • 183b65190a Clean up logging (#15919) Lianmin Zheng 2025-12-27 15:27:12 -08:00
  • 2af955e16c [model-gateway]: remove unnecessary comment (#15947) Ratish P 2025-12-27 22:33:02 +04:00
  • 0cd2b719a5 [diffusion] chore: remove useless params (#15925) Yuhao Yang 2025-12-28 01:01:08 +08:00
  • 39d56196a0 [diffusion] logging: log available gpu mem while loading and generating (#15936) Mick 2025-12-28 00:34:58 +08:00
  • 41addd2e08 chore: bump mooncake version to 0.3.8 (#15886) Shangming Cai 2025-12-27 23:34:24 +08:00
  • 5c393e8153 Fix temp_prefill_info assertion error in PP disaggregation mode (#15943) Hudson Xing 2025-12-27 22:46:55 +08:00
  • 3645ed0f73 [model-gateway] Add PrefixHash load balancing policy for KV cache-aware routing (#15935) Simo Lin 2025-12-27 05:58:05 -05:00
  • ca740a41f3 [model-gateway]: fix grpc embedding test (#15934) Ratish P 2025-12-27 14:56:54 +04:00
  • 0e25aa439b [model-gateway] optimize radix tree memory and reduce allocations (#15933) Simo Lin 2025-12-27 05:47:30 -05:00
  • 60a230b1fd [NPU] Support w4a8 with activation clip (#14736) jiaming1130 2025-12-27 16:19:46 +08:00
  • aa89c6a7e2 [diffusion] refactor: unify model loading and offloading behavior (#15923) Mick 2025-12-27 16:18:24 +08:00
  • 171912a9e3 [model-gateway] Add consistent hashing for ManualPolicy routing (#15907) Simo Lin 2025-12-27 01:53:03 -05:00
  • faecd37ed4 Add Mimo-v2-flash model to ci test (#15887) Ke Bao 2025-12-27 14:18:08 +08:00
  • 9ad546d7e8 Tiny cleanup the models' name in test_utils (#15920) Liangsheng Yin 2025-12-27 14:13:23 +08:00
  • 29ce7b3612 [diffusion] chore: remove stepvideo code (#15918) Yuhao Yang 2025-12-27 13:25:05 +08:00
  • a8380ded71 Add a test case for crash dump (#15905) Lianmin Zheng 2025-12-26 19:39:28 -08:00
  • 4edee6954a [model-gateway] add JWT/OIDC authentication for control plane APIs (#15850) Simo Lin 2025-12-26 21:30:00 -05:00