onnxruntime

mirror of https://github.com/saymrwulf/onnxruntime.git synced 2026-06-29 03:30:52 +00:00

History

Jing Fang 5dee95fa10 [CUDA] Support CUDA EP blocked quantization in Q/DQ ops. (#21846 ) ### Description 1. Added CUDA EP support for blocked quantization in QuantizeLinear and DequantizeLinear ops. 2. Currently CUDA EP blocked quantization only supports int4/uint4 quantized types and float32/float16 unquantized types. 3. Added CUDA EP support in QDQ selector/action transformer. CUDA EP is only added to DQ + MatMul -> MatMulNBits rule. Other rules' EP support are not changed. ### Motivation and Context ONNX opset 21 introduced blocked quantization for Q/DQ opts. ORT originally only supports CPU EP blocked quantization.		2024-08-30 18:28:00 -07:00
..
c_cxx	Remove extraneous javascript includes (#17558 )	2023-09-14 20:43:24 -07:00
execution_providers/images
images
python	[Fix] Make python API doc generation in Microsoft-hosted Agent (#21766 )	2024-08-20 23:32:38 +08:00
ABI_Dev_Notes.md	Fix a typo in ABI_Dev_Notes.md (#17832 )	2023-10-09 07:51:34 -07:00
Android_testing.md
C_API_Guidelines.md
cmake_guideline.md
Coding_Conventions_and_Standards.md	[docs] Specify Objective-C max line length. (#16503 )	2023-06-28 16:58:23 -07:00
ContribOperators.md	Phi3 MoE cuda kernel (#21819 )	2024-08-27 09:21:30 -07:00
FAQ.md	[Technical docs] Fixed a couple of old links in `FAQ.md` (#17415 )	2023-09-26 13:38:24 -07:00
How_To_Update_ONNX_Dev_Notes.md	Update Dockerfile.cuda (#21042 )	2024-06-13 23:50:03 -07:00
Memory_Optimizer.md	Flash attention recompute (#20603 )	2024-05-21 13:38:19 +08:00
Model_Test.md	Update docs/Model_Test.md (#11466 )	2024-05-15 11:33:11 -07:00
NotesOnThreading.md
ONNX_Runtime_Server_Usage.md
onnxruntime_dependencies.dot
onnxruntime_dependencies.png
onnxruntime_extensions.md	Remove the extensions submodule (#17097 )	2023-08-14 10:16:33 -07:00
OperatorKernels.md	[CUDA] Support CUDA EP blocked quantization in Q/DQ ops. (#21846 )	2024-08-30 18:28:00 -07:00
ORT_Format_Update_in_1.13.md
ORT_Use_Triton_Kernel.md	Rename a mispelled filename in the documentation (#21066 )	2024-06-17 18:18:41 +02:00
ORTModule_Convergence_Notes.md	Fix and enable few ORTModule Unit Tests (#19847 )	2024-03-12 10:49:19 +08:00
ORTModule_ModuleWithLoss_Wrapper.md	add steps to write modulewithloss wrapper (#16486 )	2023-07-11 09:07:35 +08:00
ORTModule_PythonOp_Notes.md	Add document for PythonOp (#17888 )	2023-10-12 08:36:22 +08:00
ORTModule_Training_Guidelines.md	Adds ATen fallback for scaled_dot_product_attention (#21107 )	2024-07-22 16:37:04 -07:00
PR_Guidelines.md
Privacy.md
Reduced_Operator_Kernel_build.md
ReleaseManagement.md
Roadmap.md
Server.md
TVM_EP.md	Fix: update hyperlinks to the Jupyter notebooks (#16145 )	2023-08-21 09:53:05 -07:00
Versioning.md
WinML_principles.md