v3.9 (#2185)
* v3.8 update x * fix blackwell gg * doc change * doc change * doc change --------- Co-authored-by: yuzhai <yuzhai@nvidia.com> Co-authored-by: Haicheng Wu <haichengw@nvidia.com> Co-authored-by: Haicheng Wu <57973641+hwu36@users.noreply.github.com>
This commit is contained in:
co-authored by
yuzhai
Haicheng Wu
Haicheng Wu
parent
8c4d1dc47d
commit
62750a2b75
@@ -530,13 +530,13 @@ If the scale factor tensor exceeds M128xSF4, it indicates that there are multipl
|
||||
<img src="../images/narrow_precison_multiple_block_sf_layout.png" alt="/narrow_precison_multiple_block_sf_layout.png"/>
|
||||
</p>
|
||||
|
||||
The creation of scale factor tensors' layouts are tedious. CUTLASS provides `Sm100BlockScaledConfig` to create these layouts easily
|
||||
The creation of scale factor tensors' layouts are tedious. CUTLASS provides `Sm1xxBlockScaledConfig` to create these layouts easily
|
||||
(See [sm100_blockscaled_layout.hpp](cutlass/include/cutlass/detail/sm100_blockscaled_layout.hpp)).
|
||||
The interface to create SFA and SFB tensor layouts is as follows:
|
||||
|
||||
```cpp
|
||||
auto problem_shape = make_shape(M, N, K, L);
|
||||
using SfConfig = Sm100BlockScaledConfig<SFVecSize>;
|
||||
using SfConfig = Sm1xxBlockScaledConfig<SFVecSize>;
|
||||
|
||||
// SFA shape: ((32,4), ceil(M/128)), ((SFVecSize,4), ceil(K/4), L)
|
||||
auto layout_sfa = SfConfig::tile_atom_to_shape_SFA(problem_shape);
|
||||
|
||||
Reference in New Issue
Block a user