laoyao0822
d21952b903
Reduce inactive NSA index-cache transfer safely
...
Centralize the IndexCache skip formula and thread the resulting active logical index layers into NSA KV pools. HiCache now skips only the indexer H2D/D2H payload for inactive target layers while preserving per-layer MLA KV transfer, keeping allocation shape unchanged for this phase.
Constraint: P0-P2 must not compact device or host allocation yet; prefill/decode state transfer still has no logical layer-id metadata.
Rejected: Recompute the skip formula separately in mem_cache | formula drift would corrupt cache or waste transfers when offset/pattern settings change.
Rejected: Skip whole-layer HiCache load/backup | MLA KV remains required for every attention layer.
Confidence: medium
Scope-risk: moderate
Directive: Before enabling compact state buffers or compact allocation, add layer-id metadata validation to PD transfer.
Tested: Local py_compile for touched files; remote pytest in g0034 container: test_nsa_index_layers.py and TestNSAIndexerPageIndices, 20 passed.
Not-tested: ETE replay/GSM8K with --nsa-index-topk-freq 4; PD state-transfer compaction remains unimplemented.
2026-06-10 04:28:26 +08:00
Xinyuan Tong
6b8a6545b2
Add Mistral Small 4 (Pixtral) support ( #20708 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Alex Nails <alexnails@radixark.ai >
Co-authored-by: Dimitrios Bariamis <12195802+dbari@users.noreply.github.com >
Co-authored-by: dbari <dbari@users.noreply.github.com >
2026-03-18 14:15:32 -07:00
Xinyuan Tong
d1e95af282
Upgrade transformers==5.3.0 ( #17784 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com >
Co-authored-by: Alison Shao <alisonshao@mac.lan >
Co-authored-by: Mick <mickjagger19@icloud.com >
2026-03-18 13:50:43 -07:00
ishandhanani
8f0f36c64b
[1/2] Add ModelExpress coordination for remote instance weight loading - matching TP ( #19920 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Ishan Dhanani <ishan@dhanani.dev >
2026-03-18 13:38:32 -07:00
shuwenn
1ac6a26464
fix: Nemotron chunk size alias ( #20458 )
2026-03-14 23:23:39 -07:00
Mohammad Miadh Angkad
75a7879fd4
[Model] Support Nemotron 3 Super NVFP4 ( #20407 )
2026-03-14 00:56:26 -07:00
R0CKSTAR
db97f193b7
[diffusion][llm] macOS support ( #19549 )
...
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com >
Co-authored-by: Mick <mickjagger19@icloud.com >
2026-03-10 13:11:07 -07:00
sjqgogogogo
eb4ba1bde2
Feature/support longcat flash lite ( #17838 )
...
Co-authored-by: sunjiaqi11 <sunjiaqi11@meituan.com >
Co-authored-by: ispobock <ispobaoke@gmail.com >
2026-03-09 23:00:11 +08:00
danielafrimi
f8bbf56de7
Refactor NemotronHConfig to canonical layers_block_type and add MTP block-type support ( #19950 )
...
Signed-off-by: dafrimi <dafrimi@nvidia.com >
2026-03-06 23:22:03 -08:00
xdtbynd
0252ca8255
[Bugfix] Fix the bug blocking the startup of Llama-3.2-11b-Vision-Instruct ( #19638 )
...
Co-authored-by: sglang-npu-bot <sglangnpu@163.com >
2026-03-06 16:21:50 +08:00
rakesh
a710b7d791
[Sarvam] Add inference support for Sarvam MoE LLMs ( #18938 )
...
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-04 15:28:00 -08:00
strgrb
738ebfd330
KDA: fuse qkv conv and support stride for fused_sigmoid_gating_delta_rule_update_kernel ( #19506 )
2026-03-04 22:45:53 +08:00
Praneth Paruchuri
f7897def96
[Feature] Improve weight loading log ( #18651 )
...
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-03 14:16:13 -08:00
Shaun Kotek
4c95953b77
Fix/nemotron mtp quantaized ( #19433 )
2026-03-03 01:07:46 -08:00
Yuwei An
c64274c746
Piecewise Cuda Graph set default ( #16331 )
2026-03-02 23:18:07 +08:00
Xinyuan Tong
581bf53e03
Whisper model support & /v1/audio/transcriptions endpoint & benchmark ( #16983 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: MahmoudAshraf97 <hassouna97.ma@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-23 17:28:37 -08:00
Shivam jindal
4f0409f8aa
[Model] Add Qwen3ForRewardModel and fix Qwen3ForSequenceClassification ( #17992 )
...
Co-authored-by: yes-its-shivam <yes-its-shivam@users.noreply.github.com >
2026-02-16 19:44:41 +08:00
SoluMilken
07a24f1a38
update pre-commit config ( #18860 )
2026-02-16 00:18:31 +08:00
andyluo7
944a9f6fcf
Fix/qwen3 5 amd rope cutedsl fallback ( #18753 )
...
Co-authored-by: seungrokj <seungrok.jung@amd.com >
2026-02-14 22:09:44 -08:00
Bhavneek Singh
1ce3420784
Model: Support IBM Granite (Dense/Mamba + MoE) ( #18040 )
2026-02-15 11:24:41 +08:00
Ke Bao
a0ebaa6498
Cleanup debug log for Ring model ( #18793 )
2026-02-13 18:36:20 +08:00
ant-yy
d97eb111a3
Support LingV2_5 model ( #18598 )
...
Co-authored-by: zhangkaihong.zkh <zhangkaihong.zkh@antgroup.com >
Co-authored-by: 有禾 <zhangdonghao.zdh@antgroup.com >
Co-authored-by: yudian0504 <138860534+yudian0504@users.noreply.github.com >
Co-authored-by: 悠扬 <youyang.zmy@antgroup.com >
Co-authored-by: xinxingyang <xinxing.yangxx@antgroup.com >
Co-authored-by: zmy460290 <zmy460290@antgroup.com >
2026-02-13 16:09:15 +08:00
Zhiyu
7e262b6496
Update modelopt quantization config parsing ( #13919 )
...
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com >
2026-02-12 11:08:29 +08:00
Jiayi Yan
539bbf485c
[Bugfix] fix config bug caused by PR #18273 ( #18535 )
2026-02-12 09:26:46 +08:00
Piotr Mazurek
ded068a76e
Add LMF2 MoE model architecture ( #17997 )
2026-02-12 01:03:43 +08:00
McZyWu
4f7422f7ba
[NPU] support model skywork-reward-gemma2-2-27B-v0.2 ( #16947 )
...
Co-authored-by: cy <chenyang08056032@163.com >
2026-02-11 15:34:53 +08:00
Zheng Li
44603764d6
fix(config): Support setting Mamba state dtype via config file ( #18532 )
...
Co-authored-by: 瑀澈 <yuche.lz@alibaba-inc.com >
2026-02-11 00:20:06 +08:00
Xinyuan Tong
398b81f78c
Support GlmMoeDsaForCausalLM ( #18521 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Signed-off-by: BBuf <1182563586@qq.com >
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com >
Co-authored-by: BBuf <1182563586@qq.com >
2026-02-10 15:20:10 +08:00
Xinyuan Tong
e8a2c13380
Deepseekv32 compatibility with transformers v5 ( #18297 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com >
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com >
2026-02-10 14:50:40 +08:00
Zheng Li
27c447653d
model: support Qwen3.5 ( #18489 )
...
Co-authored-by: 瑀澈 <yuche.lz@alibaba-inc.com >
2026-02-10 00:27:59 +08:00
Piotr Mazurek
656a3d742e
Add tensor parallelism support to LFM2 ShortConv layers ( #17777 )
2026-02-09 00:52:47 +08:00
RunningLeon
3e7ecb78a6
model: support interns1-pro ( #18145 )
...
Co-authored-by: Ke Bao <ispobaoke@gmail.com >
2026-02-05 00:22:44 +08:00
Yuhao Yang
980d2936cd
model: support Step-3.5-Flash ( #18084 )
...
Co-authored-by: ltd0924 <ltd0924@sina.com >
2026-02-03 00:40:07 +08:00
Ke Bao
d396650bd2
Fix swa kv cache memory allocation ( #18039 )
2026-02-01 14:26:51 +08:00
khalilzhk
429ef988bc
[BugFix] Fix draft model specified config file ( #17815 )
2026-01-31 20:45:36 -08:00
Kaixi
2b2515423a
Skipped warning on sm100 ( #18000 )
2026-01-31 20:21:03 -08:00
Changhun Lee
c04efe030a
[Model] Add K-EXAONE model support ( #16294 )
...
Signed-off-by: lkm2835 <lkm2835@gmail.com >
Co-authored-by: lgai-exaone <exaonemodels@lgresearch.ai >
Co-authored-by: lkm2835 <lkm2835@gmail.com >
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
2026-01-30 20:01:14 +08:00
jianan-gu
336dc4579e
[CPU] Optimize Qwen3-next model on CPU ( #12525 )
...
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com >
Co-authored-by: Fan Yin <1106310035@qq.com >
2026-01-29 22:03:58 -08:00
baonudesifeizhai
84ab611af8
model: support DeepSeek-OCR-2 ( #17897 )
2026-01-30 09:49:51 +08:00
Shivam jindal
0769de9b0f
Support LightOnOCR-2-1B ( #17806 )
2026-01-29 23:03:41 +08:00
R0CKSTAR
d3cdee0a04
[MUSA][4/N] Add common device utilities, distributed backend, and custom op wiring ( #17246 )
...
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com >
2026-01-28 23:13:24 -08:00
Yuxuan Zhang
7106f6c8e1
[GLM-OCR] Support GLM-OCR Model ( #17582 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
2026-01-26 22:24:00 -08:00
Yuhao Yang
479ab7a4e7
model: support Kimi-K2.5 ( #17789 )
...
Co-authored-by: Mick <mickjagger19@icloud.com >
2026-01-27 10:57:00 +08:00
McZyWu
2734b23481
accuracy enhancement for baichuan2-13B for npu ( #16868 )
...
Co-authored-by: cy <chenyang08056032@163.com >
2026-01-26 16:14:35 +08:00
CSWYF3634076
1a19b3987d
[Model] Add Ernie4.5 VL model support ( #15679 )
...
Signed-off-by: CSWYF3634076 <wangyafeng@baidu.com >
Signed-off-by: wangyafeng <wangyafeng@baidu.com >
2026-01-25 22:36:29 -08:00
chenxu214
444b9521e4
[Bugfix]Repeated add modelslim quant_config and bugfix with "enable-piecewise-cuda-graph" on NPU ( #17511 )
2026-01-26 09:51:07 +08:00
Ke Bao
30ece5e1d6
Fix swa memory pool size with spec ( #17630 )
2026-01-25 14:10:43 +08:00
Xinyuan Tong
37c04c2245
fix: Refactor register_image_processor to use kwarg instead of positional arg ( #17685 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
2026-01-24 15:31:01 -08:00
Lianmin Zheng
56e6652d1d
Lazy import torchao ( #17626 )
2026-01-22 22:04:51 -08:00
JiaruiChang5268
c0b5a180fe
[NPU]bugfix: fix for dsv3.2 and dsvl2 ( #17007 )
...
Co-authored-by: Hexq0210 <893781835@qq.com >
Co-authored-by: liupeng374 <782420244@qq.com >
Co-authored-by: cy <chenyang08056032@163.com >
2026-01-23 11:15:15 +08:00