This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
11
Packages
Projects
Releases
Wiki
Activity
Files
cc0485bef29831f2fcf707ecc1a371be0c7bc816
sglang
/
test
/
srt
T
History
fzyzcjy
923f518337
CUDA-graph-compatible releasing and resuming KV cache and model weight memory (
#2630
)
2025-01-13 11:38:51 -08:00
..
configs
…
models
…
sampling
/penaltylib
…
double-sparsity-config-Llama-3.1-8B-Instruct.json
…
experiment_runner.py
…
kv_cache_scales_llama3_1_8b.json
…
run_suite.py
CUDA-graph-compatible releasing and resuming KV cache and model weight memory (
#2630
)
2025-01-13 11:38:51 -08:00
test_abort.py
…
test_bench_one_batch.py
…
test_bench_serving.py
…
test_cache_report.py
…
test_chunked_prefill.py
…
test_create_kvindices.py
…
test_data_parallelism.py
…
test_double_sparsity.py
…
test_dp_attention.py
…
test_eagle_infer.py
…
test_ebnf_constrained.py
…
test_embedding_openai_server.py
…
test_eval_accuracy_large_chunked_prefill.py
…
test_eval_accuracy_large_mixed_chunked_prefill.py
…
test_eval_accuracy_large.py
…
test_eval_accuracy_mini.py
…
test_fp8_kvcache.py
…
test_fused_moe.py
…
test_get_weights_by_name.py
…
test_gguf.py
Revert "Revert "[FEAT] Support GGUF format"" (
#2287
)
2024-11-30 22:14:48 -08:00
test_input_embeddings.py
…
test_json_constrained.py
…
test_large_max_new_tokens.py
…
test_matched_stop.py
…
test_metrics.py
…
test_mla_fp8.py
…
test_mla.py
…
test_models_from_modelscope.py
…
test_moe_ep.py
…
test_moe_eval_accuracy_large.py
Fix linear.py and improve weight loading (
#2851
)
2025-01-13 01:39:14 -08:00
test_nightly_gsm8k_eval.py
…
test_nightly_human_eval.py
Fix nightly accuracy tests (
#2780
)
2025-01-07 21:02:35 -08:00
test_nightly_math_eval.py
…
test_no_chunked_prefill.py
…
test_no_overlap_scheduler.py
…
test_openai_server.py
…
test_pytorch_sampling_backend.py
…
test_radix_attention.py
…
test_release_memory_occupation.py
CUDA-graph-compatible releasing and resuming KV cache and model weight memory (
#2630
)
2025-01-13 11:38:51 -08:00
test_retract_decode.py
…
test_schedule_policy.py
…
test_server_args.py
…
test_session_control.py
…
test_skip_tokenizer_init.py
…
test_srt_endpoint.py
…
test_srt_engine_with_quant_args.py
…
test_srt_engine.py
…
test_torch_compile_moe.py
Improve torch compile for fused moe (
#2327
)
2024-12-03 01:58:25 -08:00
test_torch_compile.py
…
test_torch_native_attention_backend.py
Add a simple torch native attention backend (
#2241
)
2024-12-01 03:01:25 -08:00
test_torch_tp.py
…
test_torchao.py
…
test_triton_attention_backend.py
…
test_triton_attention_kernels.py
…
test_update_weights_from_disk.py
…
test_update_weights_from_distributed.py
…
test_update_weights_from_tensor.py
…
test_vision_chunked_prefill.py
[feat] Enable chunked prefill for llava-onevision (
#2412
)
2024-12-09 09:52:38 -08:00
test_vision_openai_server.py
…