tpoisonooo / how-to-optimize-gemm Goto Github PK

row-major matmul optimization

Home Page: https://zhuanlan.zhihu.com/p/65436463

License: GNU General Public License v3.0

C 10.65% MATLAB 0.04% Makefile 0.06% Shell 0.01% Python 0.05% C++ 88.18% Assembly 0.01% Cuda 0.59% Objective-C 0.20% Mathematica 0.18% M 0.01% CMake 0.02%

gemm-optimization armv7 arm64 cuda cuda-kernel ptx vulkan int4

how-to-optimize-gemm's Issues

cuda代码中错误

cuda代码中，MMult_cuda_7.cu中30行，b_ptr += 64 * k，应该是b_ptr += 64 * n，因为是方阵，所以结果对上了

about ldg32_nc_0

https://github.com/tpoisonooo/how-to-optimize-gemm/blob/master/cuda/MMult_cuda_12.cu: 20,21
I'm a beginner of CUDA&&PTX, I want to know what does these two PTX use for?
"{.reg .pred p;\n"
"mov.b32 %0, 0;\n"
is it useless code?

Why 2 kernels in armv8/MMult_4x4_13?

Is that for performance reason or just comparison?

how to overlap the share2register and computing process?

I have another question about MMult_cuda_12.cu
Honestly, I don't know how to overlap the share2register and computing process. Is it the asm(PTX) that make them run parallelly? The instructions are sequantially, so how could these two parts of code hide each other?
part1: loading shared-memory to panel
lds128(panelA[pp][0], panelA[pp][1], panelA[pp][2], panelA[pp][3],
aptr_base + ((subk + 1) % 8) * SMEM_LDA * sizeof(float));
lds128(panelA[pp][4], panelA[pp][5], panelA[pp][6], panelA[pp][7],
aptr_base + (((subk + 1) % 8) * SMEM_LDA + 64) * sizeof(float));
lds128(panelB[pp][0], panelB[pp][1], panelB[pp][2], panelB[pp][3],
bptr_base + ((subk + 1) % 8) * SMEM_LDB * sizeof(float));
lds128(panelB[pp][4], panelB[pp][5], panelB[pp][6], panelB[pp][7],
bptr_base + (((subk + 1) % 8) * SMEM_LDB + 64) * sizeof(float));

part2: computing the result of panel-data
#pragma unroll
for (int i = 0; i < 8; ++i) {
#pragma unroll
for (int j = 0; j < 8; ++j) {
sum[i][j] += panelA[subk % 2][i] * panelB[subk % 2][j];
}
}

v11, line 20: why we define a p but not used it? line 29, why we use long type?

Hi! I am learning your wonderful code but get puzzled by some assembly code. Excuse me that I am a new assembly learner. I searched for some documents and still find these two points strange:
v11, line 20: why we define a p but not used it?
line 29, why we use long type?

Thank you!!!

packing analysis

运行make run （MMult_cuda_2)时报错 CUDA error

T4卡上运行时出现报错 CUDA error at test_MMult.cpp:110 code=999 "cudaEventSynchronize(stop)",想问一下是什么原因呢

question about gflops benchmark

gflops benchmark中的 #define OP_FLOATS (80), 这里的80是怎么计算的呢？

cuda版本非m=n=k运算出错

如kernel_v3中:
float *begin_a = a + by * BLOCK * k; //by->n
float *begin_b = b + bx * BLOCK; //bx->m

当A,B不为方阵时会出错,例如m=k=256,n=128.

fp32 gemm存在微小误差

A(i, j) = (2.0 * (float)drand48() - 1.0)
误差为
4.577637e-05

如果数字更大，误差更大

关于测试浮点峰值的问题

我现在跑的芯片型号是NVIDIA，ARMv8 Processor rev 0 (v8l)。
我看知乎文章里说测试浮点峰值时FMA指令的排布数量 = FMA的发射数 * FMA指令的延迟。我并没有查到上面这个芯片的手册。但是我看了A57的手册，里面是这样记录的：

FMA指令的延迟是10，吞吐量是2。我不太清楚这个吞吐是否代表着芯片可以同时发射两条FMA指令（是芯片发射吗），但是我分别放置了10条FMA指令（OP_FLOATS = 80）和20条FMA指令（OP_FLOATS = 160）都测试了，发现在10条的时候是16.095492 GFLOPS， 20条是 18.759214 GFLOPS。这是什么原因呢？
我的猜测有两个：
1.10条FMA指令确实不是测试这款芯片的浮点峰值所需要的指令数。
2.可能编译器自动开启了多线程？这个比较有可能，因为从4条指令到10条指令性能差不多翻倍，但是10-20只增加了一点。

Row-majored or Column-majored

I found the definition of MMult0.c and MMult1.c for multi-dimensional array storage, would you like to take a look and see whether both of them are correct?
For MMult0.c

how-to-optimize-gemm/src/HowToOptimizeGemm/MMult0.c

Lines 3 to 5 in ddab74e

 #define A(i,j) a[ (i)*lda + (j) ] 

 #define B(i,j) b[ (i)*ldb + (j) ] 

 #define C(i,j) c[ (i)*ldc + (j) ]

For MMult1.c

how-to-optimize-gemm/src/HowToOptimizeGemm/MMult1.c

Lines 3 to 5 in ddab74e

 #define A(i,j) a[ (j)*lda + (i) ] 

 #define B(i,j) b[ (j)*ldb + (i) ] 

 #define C(i,j) c[ (j)*ldc + (i) ]

As you see, the positions of i and j are inverted.

In cuda v12, the a_sts_addr and a_base are redundant?

Hi! I think the a_sts_addr and a_base are the same? So maybe we can delete one?

Thank you!!!

tpoisonooo / how-to-optimize-gemm Goto Github PK

how-to-optimize-gemm's Issues

cuda代码中错误

about ldg32_nc_0

Why 2 kernels in armv8/MMult_4x4_13?

how to overlap the share2register and computing process?

v11, line 20: why we define a p but not used it? line 29, why we use long type?

packing analysis

运行make run （MMult_cuda_2)时报错 CUDA error

question about gflops benchmark

cuda版本非m=n=k运算出错

fp32 gemm存在微小误差

关于测试浮点峰值的问题

Row-majored or Column-majored

In cuda v12, the a_sts_addr and a_base are redundant?

Recommend Projects

React

Vue.js

Typescript

TensorFlow

Django

Laravel

D3

Recommend Topics

javascript

web

server

Machine learning

Visualization

Game

Recommend Org

Facebook

Microsoft

Google

Alibaba

D3

Tencent

	#define A(i,j) a[ (i)*lda + (j) ]
	#define B(i,j) b[ (i)*ldb + (j) ]
	#define C(i,j) c[ (i)*ldc + (j) ]

	#define A(i,j) a[ (j)*lda + (i) ]
	#define B(i,j) b[ (j)*ldb + (i) ]
	#define C(i,j) c[ (j)*ldc + (i) ]