Update link to E2E notebook in LLaMA-2 blog (#20724)

### Description
This PR updates a reference link in the LLaMA-2 blog post and fixes a
word formatting issue.

### Motivation and Context
With these changes, the link to the example E2E notebook works again.
This commit is contained in:
kunal-vaishnavi 2024-05-19 22:51:46 -07:00 committed by GitHub
parent ca6b0f8cb2
commit 5fd617acc0
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -102,7 +102,7 @@
>batch size * (prompt length + token generation length) / wall-clock latency</i
> where wall-clock latency = the latency from running end-to-end and token generation length =
256 generated tokens. The E2E throughput is 2.4X more (13B) and 1.8X more (7B) when compared to
PyTorch compile. For higher batch size, sequence length like 16, 2048 pytorch eager times out,
PyTorch compile. For higher batch size, sequence length pairs such as (16, 2048), PyTorch eager times out,
while ORT shows better performance than compile mode.
</p>
<div class="grid grid-cols-1 lg:grid-cols-2 gap-4">
@ -151,7 +151,7 @@
<p class="mb-4">
More details on these metrics can be found <a
href="https://github.com/microsoft/onnxruntime-inference-examples/blob/main/python/models/llama2/README.md"
href="https://github.com/microsoft/onnxruntime-inference-examples/blob/main/python/models/llama/README.md"
class="text-blue-500">here</a
>.
</p>