Commit Graph
100 Commits
Author SHA1 Message Date
Kangyan-Zhou 62361e8fd5 Temporarily adjust the scheulde for pr-test.yml to 12 hours (#20951) 2026-03-19 14:48:10 -07:00
Kangyan-ZhouandClaude Opus 4.6 b6055e59cd [HiCache] Reduce per-request backup log noise (#20813)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 22:47:14 -07:00
Kangyan-Zhou 3d8fc9a0ca Revert "[Nvidia] Add trtllm mnnvl allreduce with unified flashinfer allreduce fusion api" (#20792) 2026-03-17 11:59:02 -07:00
Kangyan-ZhouandClaude Opus 4.6 f5a4a5429f Revert early HTTP port reservation (#17754, #19805) (#20468)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 16:17:33 -07:00
Kangyan-ZhouandClaude Opus 4.6 7a12255b6e fix: set first_token_time before computing decode_throughput for single-batch completions (#19984)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 16:11:41 -08:00
Kangyan-ZhouandClaude Opus 4.6 e89069ee64 Fallback to torch.cuda.mem_get_info() when nvidia-smi is unavailable (#18957)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 15:00:08 -08:00
Kangyan-ZhouandClaude Opus 4.6 25e678d933 [diffusion] endpoint: add /server_info and /model_info endpoints for gateway discovery (#20020)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 21:36:13 +08:00
Kangyan-Zhouand 198381d9ce Add SSL/TLS support for HTTP and gRPC servers (#18973)
Co-authored-by: guys@spotify.com
2026-03-04 19:27:16 -08:00
Kangyan-ZhouandClaude Opus 4.6 0eb64c1e72 [smg] Extract tokenizer_path from /model_info into discovered labels (#19905)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 17:12:20 -08:00
Kangyan-ZhouandClaude Opus 4.6 c03deb8175 Fix disagg PD bootstrap and KV transfer metrics (#19009)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 09:08:10 -08:00
Kangyan-ZhouandClaude Opus 4.6 dc92f88a21 Enhance bench_multiturn.py with OpenAI API support and richer metrics (#19724)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 13:48:04 -08:00
Kangyan-ZhouandClaude Opus 4.6 ec9775491f Add bisect ci claude code skill (#19649)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 20:53:40 -08:00
Kangyan-ZhouandClaude Opus 4.6 98224de29b [Bugfix] Add missing auto_create_handle_loop to communicator methods (#19610)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 15:00:05 -08:00
Kangyan-ZhouandClaude Opus 4.6 dc02e5bea7 [HiCache] Re-land spec v2 + decode KV cache offloading compatibility (#19615)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 13:58:31 -08:00
Kangyan-Zhou dcf462cfba Revert "[HiCache] Enable spec v2 + decode KV cache offloading compatibility" (#19613) 2026-02-28 21:54:32 -08:00
Kangyan-ZhouandClaude Opus 4.6 8167346609 [HiCache] Enable spec v2 + decode KV cache offloading compatibility (#19518)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 21:53:52 -08:00
Kangyan-Zhou 306c552639 Revert "Fix HybridAttnBackend forward for linear attention" (#19356) 2026-02-25 11:49:50 -08:00
Kangyan-ZhouandClaude Opus 4.6 8aeb16f3fc fix: add missing blank line after docstring in serving_transcription.py (#19206)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 18:29:38 -08:00
Kangyan-ZhouandClaude Opus 4.6 b2d8cc8cf0 Fix dev Docker build OOM on ARM64 cu13 by adding docker system prune (#18947)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 10:29:31 +08:00
Kangyan-ZhouandClaude Opus 4.6 ae95869292 Enable SGLANG_ENABLE_SPEC_V2 for nightly speculative decoding tests (#18719)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-14 23:00:33 +08:00
Kangyan-Zhou 3a1c388b43 Update performance dashboard for nightly tests (#18824) 2026-02-14 09:28:28 +08:00
Kangyan-ZhouandClaude Opus 4.6 eccf875d49 [CI] Revive 8-GPU trace upload in nightly test workflow (#18820)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-14 08:37:08 +08:00
Kangyan-Zhou 710d873ba6 Update notified user in post_ci_failures_to_slack.py (#18817) 2026-02-14 06:48:56 +08:00
Kangyan-Zhou 1b8f68af57 Fix B200 installation issue (#18725) 2026-02-12 22:06:23 +08:00
Kangyan-Zhou f116b3a51b Make PR based docker and pypi workflow work for forked PR (#18720) 2026-02-12 21:05:17 +08:00
Kangyan-Zhou 0db6fd4dbe Revert broken sgl_kernel exclusion patterns in paths-filter (#18193) 2026-02-03 10:56:44 -08:00
Kangyan-Zhou cd31540fd7 Improve Per Commit Test job filtering for sglang-kernel (#18054) 2026-02-01 21:15:23 -08:00
Kangyan-Zhou 9c168fcac7 Fix Diffusion Request Validation to allow missing input artifacts if the input only contains text (#16610) 2026-01-31 23:38:40 -08:00
Kangyan-Zhou e5ac6229e1 Fix installation script for H200 runners (#18050) 2026-01-31 23:30:51 -08:00
Kangyan-Zhou e884b17632 Fix rerun stage command with merged commit history (#17960) 2026-01-31 20:37:55 -08:00
Kangyan-Zhou d443d2d2ae Improve error output in tnightly tets (#18053) 2026-01-31 19:26:19 -08:00
Kangyan-Zhou a16aba8340 Increase install dependency for gb200 (#17977) 2026-01-29 23:55:25 -08:00
Kangyan-ZhouandClaude Opus 4.5 606ff09ef8 [Fix] Remove unused Type import in gpt_j.py (#17975)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-29 21:11:11 -08:00
Kangyan-ZhouandClaude Opus 4.5 2cd2c3118d Add concurrency tracking to runner utilization report (#17963)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-29 17:31:55 -08:00
Kangyan-Zhou c3bf53c7c1 Fix ci weight validation logic to check the safetensor completeness (#17917) 2026-01-29 13:00:42 -08:00
Kangyan-Zhou c0b4dd68a2 Add a performance dashboard server and frontend for nightly CUDA tests (#17725) 2026-01-27 22:22:33 -08:00
Kangyan-Zhou 48f4340b14 Exclude some diffusion package for ARM in docker release (#17745) 2026-01-25 23:32:39 -08:00
Kangyan-Zhou 52e0f65fce Update nightly-test-nvidia.yml to remove push trigger (#17625) 2026-01-25 21:42:36 -08:00
Kangyan-Zhou 5aaedac3c6 Add EP=2 to qwen235b nightly tests (#17738) 2026-01-25 21:36:04 -08:00
Kangyan-ZhouandXinyuan Tong 592603d77b Fix flaky streaming logprobs test by handling detokenizer text buffering (#17687)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-01-25 15:09:06 -08:00
Kangyan-Zhou 344eeaee90 Upload nightly test metrics to GH artifacts (#17696) 2026-01-25 14:35:14 -08:00
Kangyan-Zhou 7ca8c12e0e Extend b200 kernel tests timeout for CPU differences (#17718) 2026-01-25 12:59:39 -08:00
Kangyan-Zhou 8d3e1ac0c8 Add an all type in pyproject.tml to include diffusion support (#17697) 2026-01-25 12:52:13 -08:00
Kangyan-Zhou 9123491430 A few updates to the night tests (#17694) 2026-01-25 11:20:17 -08:00
Kangyan-Zhou b829b797ef Fix slash command handler trigger condition by trimming the comments (#17691) 2026-01-24 18:54:46 -08:00
Kangyan-Zhou 69a7a70e47 Temporarily disable lora overlap loading test due to flakiness (#17683) 2026-01-24 12:24:44 -08:00
Kangyan-Zhou 137eb5b95c Fix NSA indexer test and move it to pre commit test (#17682) 2026-01-24 12:06:18 -08:00
Kangyan-Zhou 8656a146a6 Fix test timeout issue in pr-test (#17681) 2026-01-24 11:46:08 -08:00
Kangyan-Zhou f2ae066a6b Update release-branch-cut.yml for actions: write (#17539) 2026-01-21 18:03:27 -08:00
Kangyan-Zhou be5121b452 Fix NSA indexer in the nightly test (#17452) 2026-01-20 19:58:13 -08:00
Kangyan-Zhou 18e2ef09d7 Add v1/models endpoint to diffusion model APIs so that they can be discovered by model gateway (#16425) 2026-01-06 14:28:24 -08:00
Kangyan-Zhou d93f37a625 Add harvenstar to CI_PERMISSIONS.json 2026-01-04 22:35:26 -08:00
Kangyan-Zhou ca80c19b55 Revert "[diffusion] feat: support warmup with resolutions" (#16433) 2026-01-04 18:44:05 -08:00
Kangyan-Zhou 130c6911f8 Add a GH action for cherrypick (#16243) 2025-12-31 20:05:38 -08:00
Kangyan-Zhou 12b89e51d8 Add P90/99 e2e latency in bench_serving script (#16245) 2025-12-31 15:40:33 -08:00
Kangyan-Zhou 2da49eec50 Fix XEON docker release workflow (#16107) 2025-12-31 09:47:01 -08:00
Kangyan-Zhou d65ae0ec7a Use a different concurrency group for release branch testing (#16202) 2025-12-31 09:46:43 -08:00
Kangyan-Zhou fc643ffbc9 Download missing shards in model weights files when not in CI (#16211) 2025-12-31 20:42:34 +08:00
Kangyan-Zhou 0d003e34b0 Reduce CI failure monitor to run once every 12 hours (#16123) 2025-12-29 19:44:25 -08:00
Kangyan-Zhou 9c4eb46099 Add a new branch cut GH workflow, and adopt setuptools-scm for version control (#15985) 2025-12-29 13:51:21 -08:00
Kangyan-Zhou a91e072f33 Use allow auto truncate in the OpenAI API endpoint (#15369) 2025-12-25 18:45:57 -08:00
Kangyan-Zhou 011d8d8970 Reserve more memory for DeepSeekOCR model and adjust server start timeout for DeepGEMM to reduce flakiness (#15277) 2025-12-17 13:13:05 -08:00
Kangyan-Zhou 272c5fe43e Increase timeout for TestDeepseekV3MTP for potential DeepGEMM cold start (#15239) 2025-12-16 00:26:10 -08:00
Kangyan-Zhou e7157c9b77 Add cache for flashinfer installation (#15153) 2025-12-15 10:37:34 -08:00
Kangyan-Zhou 037c3982af Fix H200 CI by commenting out Warmup Weights and JIT Compilation (#15139) 2025-12-14 18:04:28 -08:00
Kangyan-Zhou 291396544e Add a special label for b200 CI runner that can run kernel tests (#15033) 2025-12-12 19:41:36 -08:00
Kangyan-Zhou b243154614 Fix CI by reverting incorrect metric check logic (#15004) 2025-12-12 10:07:38 -08:00
Kangyan-Zhou aa3716b29d Update CI_PERMISSIONS.json 2025-12-11 15:55:50 -08:00
Kangyan-Zhou c5f1e86117 Update CI_PERMISSIONS.json (#14917) 2025-12-11 10:14:42 -08:00
Kangyan-Zhou 2856624156 Only count limitations for previous runs that reaches the test stages (#14856) 2025-12-10 23:17:34 -08:00
Kangyan-Zhou 32829b1638 Remove myself to test CI gate issue (#14871) 2025-12-10 22:37:01 -08:00
Kangyan-Zhou d9dca28247 Update pr-test.yml to fix unknown job name deepep-8-gpu 2025-12-01 10:02:29 -08:00
Kangyan-Zhou 41b7aab848 Disable Deepep 8 GPU tests (#14152) 2025-12-01 09:37:38 -08:00
Kangyan-Zhou 5ddd2f6b6b Always run model evaluation even if the trace upload step fails (#14157) 2025-11-29 20:13:59 -08:00
Kangyan-Zhou 1d3d8b3418 Fix Minimax M2 loading issue (#13956) 2025-11-29 17:07:19 -05:00
Kangyan-Zhou d7cb08c5be Always run all stages in cron based PR tests (#14151) 2025-11-29 12:32:06 -08:00
Kangyan-Zhou d6c88d519d Trigger PR test on main every 3 hours instead of push event (#14130) 2025-11-29 10:12:50 -08:00
Kangyan-Zhou a102a0507a Disable Deepep 2 GPU tests (#14111) 2025-11-28 13:56:34 -08:00
Kangyan-Zhou 4c9f7c97d3 Temporarily disabled test (#14069) 2025-11-27 14:58:46 -08:00
Kangyan-Zhou 779cbc6e4b Fix Nvidia nightly test trigger params when it is triggered by parent workflow (#13966) 2025-11-26 12:01:48 -08:00
Kangyan-Zhou 13e5beeab4 Fix Deepseek v3.1 loading issue (#13954) 2025-11-25 18:28:52 -08:00
Kangyan-Zhou 03a26557b7 Fix nightly-test-nvidia.yml to have the correct trigger (#13950) 2025-11-25 17:34:34 -08:00
Kangyan-Zhou e99ca6ac74 Improve nightly tests (#13903) 2025-11-25 12:04:49 -08:00
Kangyan-Zhou 59b4d7f8d6 Fix B200 Nightly tests and move one manual test back to unit test to prevent the same issue (#13746) 2025-11-21 17:41:12 -08:00
Kangyan-Zhou cf1f0166b6 Add Qwen/Qwen1.5-MoE-A2.7B to model list (#13543) 2025-11-18 14:43:38 -08:00
Kangyan-Zhou c0d1a3383b Remove jet-ai/Jet-Nemotron-2B in nightly text tests as this is constantly failing (#13540) 2025-11-18 13:58:16 -08:00
Kangyan-Zhou e2c9a59023 Update pr-test.yml to fix invalid job name error 2025-11-17 17:08:02 -08:00
Kangyan-Zhou a1e37b0258 Temporarily comment out multimodal gen test to recover runners (#13463) 2025-11-17 16:55:04 -08:00
Kangyan-Zhou ea89a3a0c5 Fixes validation errors for Wan-AI models which store model weights in subdirectories (#13461) 2025-11-17 15:33:02 -08:00
Kangyan-Zhou 2bc7c5ebef Fix 8-gpu B200 nightly tests (#13457) 2025-11-17 13:26:56 -08:00
Kangyan-Zhou 58f8f4e408 Add missing models (#13456) 2025-11-17 13:10:37 -08:00
Kangyan-Zhou 4e19c1d541 Add missing models (#13369) 2025-11-16 00:04:47 -08:00
Kangyan-Zhou be353ffd13 Add missing models (#13351) 2025-11-15 13:51:20 -08:00
Kangyan-Zhou 8f4e18a294 Add missing model for 2-gpu-runner in nightly tests (#13311) 2025-11-14 17:34:43 -08:00
Kangyan-Zhou e9681444bd Add missing model in model validate list (#13310) 2025-11-14 17:29:21 -08:00
Kangyan-Zhou b223669136 Remove nightly b200 tests and revert a change for test file (#13305) 2025-11-14 16:30:22 -08:00
Kangyan-Zhou 49141df94a Extend lint test to test/ directory (#13247) 2025-11-14 00:01:48 -08:00
Kangyan-Zhou 04848ba7cb Add 3 models to 2 gpu runner in model downloading from nightly tests (#13261) 2025-11-13 22:59:18 -08:00
Kangyan-Zhou 9bc6a9adbe Update model weight validation logic to handle special weight file naming (#13256) 2025-11-13 22:07:49 -08:00
Kangyan-Zhou 5c2d72ba63 Bump actions/download-artifact from v4 to v6 for B200 workers (#13220) 2025-11-13 10:23:58 -08:00