onnxruntime/js/web/lib/wasm/jsep/webgpu/ops
Jiajia Qin 891fba3b9c
[js/webgpu] Optimize Gather op (#17625)
### Description
This PR optimizes the gather op, which is improved ~6ms in segment
anything model in ADL.
The problem in original algorithm is that it includes a for loop to
calculate a block size of data. However, the block size may be very
large, like `65536`. In GPU shader, we should try to avoid large loop in
shader and try to use more threads to do it parallelly.

Before:
```
[profiling] kernel "41771992|[Gather] 41771992" input[0]: [4,65536] | float32, input[1]: [1] | int64, output[0]: [1,65536] | float32, execution time: 6886207 ns
```
After:
```
[profiling] kernel "41771992|[Gather] 41771992" input[0]: [4,65536] | float32, input[1]: [1] | int64, output[0]: [1,65536] | float32, execution time: 11719 ns
2023-09-21 21:00:36 -07:00
..
3rd-party
argminmax.ts
binary-op.ts
common.ts
concat.ts
conv-grouped.ts
conv-transpose.ts
conv.ts
conv2d-mm.ts
einsum.ts
expand.ts
fuse-utils.ts
gather-elements.ts
gather.ts
gemm.ts
instance-norm.ts
layer-norm.ts
matmul.ts
pad.ts
pool.ts
reduce.ts
resize.ts
skip-layer-norm.ts
slice.ts
softmax.ts
split.ts
tile.ts
transpose.ts
unary-op.ts