onnxruntime/onnxruntime/core
Adrian Lizarraga b47e1e64d7
[QNN EP] Make offloading graph input/output quantization (to CPU) the default (#23368)
### Description
Makes the QNN provider option `offload_graph_io_quantization` enabled by
default. It was previously disabled by default.



### Motivation and Context
Enabling this option significantly decreases inference latency for many
models.
2025-02-04 11:42:46 -08:00
..
common remove log spam from cpuinfo (#23548) 2025-01-31 18:16:24 -08:00
dll fix webgpu delay load test (#23157) 2024-12-20 13:37:12 -08:00
dlpack
eager
flatbuffers
framework Fix the issue that the new generated EP context model not able to find external data (#23537) 2025-01-29 22:01:13 -08:00
graph Update BiasGelu fusion and related ops (#23518) 2025-01-30 22:53:59 -08:00
mickey
mlas [ARM CPU] hgemm optimized for gqa (#23107) 2025-01-24 15:25:24 -08:00
optimizer Update BiasGelu fusion and related ops (#23518) 2025-01-30 22:53:59 -08:00
platform [QNN EP] Make QNN EP a shared library (#23120) 2025-01-22 12:11:00 -08:00
providers [QNN EP] Make offloading graph input/output quantization (to CPU) the default (#23368) 2025-02-04 11:42:46 -08:00
quantization
session [onnxruntime/build] Add new flag enable_generic_interface to build primary EPs by default (#23342) 2025-01-28 15:24:09 -08:00
util Address CodeQL security issues on comparison of different types (#23276) 2025-01-07 17:30:44 -08:00