..
aarch32
aarch64
[aarch64] Add Sbgemm kernel to accelerate fp32 tensor matmul with bfloat16 ( #17031 )
2024-01-22 14:43:06 -08:00
amd64
Add vectorized AVX512F kernel for ReduceMaximumF32Kernel ( #20268 )
2024-04-16 13:52:43 -07:00
arm
arm64
arm64ec
i386
intrinsics
loongarch64
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
power
[MLAS] Use C-style casting for power vector instructions ( #20957 )
2024-06-06 15:11:59 -07:00
scalar
wasm_simd
Fix a bug in WASM's GEMM ( #20023 )
2024-03-23 08:53:50 -07:00
x86
x86_64
Add vectorized AVX512F kernel for ReduceMaximumF32Kernel ( #20268 )
2024-04-16 13:52:43 -07:00
activate.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
activate_fp16.cpp
amx_common.h
Change "#ifdef WIN32" to "#ifdef _WIN32" ( #19254 )
2024-01-24 14:35:44 -08:00
compute.cpp
Optimize MlasComputeSoftmax with prefetch ( #20393 )
2024-04-25 08:28:59 -07:00
convolve.cpp
convsym.cpp
dgemm.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
dwconv.cpp
erf.cpp
flashattn.cpp
Implement FlashAttention for CPU ( #20805 )
2024-07-11 14:19:59 -07:00
fp16_common.h
halfgemm.cpp
halfgemm.h
halfgemm_kernel_neon.cpp
logistic.cpp
mlasi.h
[CPU EP] Int4 support for QuantizeLinear, DequantizeLinear, and Transpose ( #20362 )
2024-05-30 18:56:24 -07:00
platform.cpp
[CPU EP] Int4 support for QuantizeLinear, DequantizeLinear, and Transpose ( #20362 )
2024-05-30 18:56:24 -07:00
pooling.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
pooling_fp16.cpp
q4_dq.cpp
[MLAS] add q4 quantize and transpose kernel to support MatMulNBits QDQ fuse ( #21054 )
2024-06-19 17:15:45 -07:00
q4_dq_cli.cpp
q4common.h
q4gemm.cpp
q4gemm.h
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
q4gemm_avx512.cpp
qdwconv.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
qdwconv_kernelsize.cpp
qgemm.cpp
qgemm.h
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
qgemm_kernel_amx.cpp
qgemm_kernel_avx2.cpp
qgemm_kernel_default.cpp
qgemm_kernel_lsx.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
qgemm_kernel_neon.cpp
qgemm_kernel_sdot.cpp
qgemm_kernel_smmla.cpp
[aarch64] Implement QGEMM kernels with UMMLA/SMMLA instructions ( #17160 )
2023-10-24 07:49:04 +10:00
qgemm_kernel_sse.cpp
qgemm_kernel_sse41.cpp
qgemm_kernel_udot.cpp
qgemm_kernel_ummla.cpp
[aarch64] Implement QGEMM kernels with UMMLA/SMMLA instructions ( #17160 )
2023-10-24 07:49:04 +10:00
qgemm_kernel_wasmsimd.cpp
qladd.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
qladd.h
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
qlgavgpool.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
qlmul.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
qpostprocessor.cpp
quantize.cpp
[CPU EP] Int4 support for QuantizeLinear, DequantizeLinear, and Transpose ( #20362 )
2024-05-30 18:56:24 -07:00
reorder.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
sbgemm.h
[aarch64] Add Sbgemm kernel to accelerate fp32 tensor matmul with bfloat16 ( #17031 )
2024-01-22 14:43:06 -08:00
sbgemm_kernel_neon.cpp
[aarch64] Add Sbgemm kernel to accelerate fp32 tensor matmul with bfloat16 ( #17031 )
2024-01-22 14:43:06 -08:00
sgemm.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
snchwc.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00
sqnbitgemm.cpp
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm.h
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_kernel_avx2.cpp
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_kernel_avx512.cpp
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_kernel_avx512vnni.cpp
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_kernel_avx_common.h
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_kernel_avx_common_fp32.h
Mlas Gemm 4bit avx2, avx512, and avx512vnni kernels ( #20163 )
2024-04-25 21:30:50 -07:00
sqnbitgemm_kernel_avx_common_int8.h
SQNBitGemm - move workspace size calculation functions to hardware-specific implementations ( #20757 )
2024-05-22 15:12:17 -07:00
sqnbitgemm_kernel_neon.cpp
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_kernel_neon.h
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_kernel_neon_fp32.cpp
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_kernel_neon_int8.cpp
[MLAS] AArch64 SQNBitGemm CompInt8 initial multi-row implementation ( #21193 )
2024-07-10 15:39:26 -07:00
sqnbitgemm_q8_block.h
SQNBitGemm - move workspace size calculation functions to hardware-specific implementations ( #20757 )
2024-05-22 15:12:17 -07:00
tanh.cpp
threading.cpp
Augment blockwise quantization ( #18101 )
2023-10-30 09:14:37 -07:00
transpose.cpp
[mlas] add loongarch lsx and lasx optimize code ( #17937 )
2023-12-07 11:15:59 -08:00