Update AVX2 docs

This commit is contained in:
Henry de Valence 2017-11-30 15:24:26 -08:00
parent 1103c5c43a
commit bf5e3581d2
2 changed files with 22 additions and 12 deletions

View file

@ -101,8 +101,8 @@ impl FieldElement32x4 {
return out; return out;
} }
// Negate variables in lanes where mask is set /// Negate variables in lanes where mask is set
// XXX fix up api /// XXX fix up api
pub fn mask_negate(&mut self, mask: u8) { pub fn mask_negate(&mut self, mask: u8) {
unsafe { unsafe {
use stdsimd::vendor::_mm256_blend_epi32; use stdsimd::vendor::_mm256_blend_epi32;
@ -114,7 +114,7 @@ impl FieldElement32x4 {
self.reduce32(); self.reduce32();
} }
// Given `self = (A,B,C,D)`, set `self = (A,B,D,C)` /// Given `self = (A,B,C,D)`, set `self = (A,B,D,C)`
pub fn swap_CD(&mut self) { pub fn swap_CD(&mut self) {
unsafe { unsafe {
use stdsimd::vendor::_mm256_shuffle_epi32; use stdsimd::vendor::_mm256_shuffle_epi32;
@ -126,7 +126,7 @@ impl FieldElement32x4 {
} }
} }
// Given `self = (A,B,C,D)`, set `self = (B - A, B + A, D - C, D + C)`. /// Given `self = (A,B,C,D)`, set `self = (B - A, B + A, D - C, D + C)`.
pub fn diff_sum(&mut self) { pub fn diff_sum(&mut self) {
/// (v0 v1 v2 v3 v4 v5 v6 v7) -> (v1 v0 v3 v2 v5 v4 v7 v6) /// (v0 v1 v2 v3 v4 v5 v6 v7) -> (v1 v0 v3 v2 v5 v4 v7 v6)
#[inline(always)] #[inline(always)]

View file

@ -12,7 +12,7 @@
//! Curve25519, using AVX2 to implement the 4-way parallel formulas of //! Curve25519, using AVX2 to implement the 4-way parallel formulas of
//! Hisil, Wong, Carter, and Dawson (HWCD). //! Hisil, Wong, Carter, and Dawson (HWCD).
//! //!
//! Their 2008 paper _Twisted Edwards Curves Revisited_, which //! Their 2008 paper [_Twisted Edwards Curves Revisited_][hwcd08], which
//! introduced the extended coordinates used in other parts of `-dalek`, //! introduced the extended coordinates used in other parts of `-dalek`,
//! also describes 4-way parallel formulas for point addition and //! also describes 4-way parallel formulas for point addition and
//! doubling: //! doubling:
@ -101,12 +101,15 @@
//! The 4-wide formulas of the HWCD paper do not seem to have been //! The 4-wide formulas of the HWCD paper do not seem to have been
//! implemented using SIMD before. The HWCD paper also describes and //! implemented using SIMD before. The HWCD paper also describes and
//! analyzes a 2-wide variant of the Montgomery ladder; this strategy was //! analyzes a 2-wide variant of the Montgomery ladder; this strategy was
//! used by Tung Chou's `sandy2x` implementation, which used a 2-wide //! used in 2015 by Tung Chou's `sandy2x` implementation, which used a 2-wide
//! field implementation in 128-bit registers. Curiously, however, //! field implementation in 128-bit vector registers. Curiously, however,
//! although the `sandy2x` paper cites the HWCD paper for extended //! although the [`sandy2x` paper][sandy2x] cites the HWCD paper for extended
//! twisted Edwards coordinates, it does not mention the 4-wide HWCD //! twisted Edwards coordinates, it does not mention the 4-wide HWCD
//! Edwards formulas or that the 2-wide Montgomery formulas it uses were //! Edwards formulas or that the 2-wide Montgomery formulas it uses were
//! previously published there. //! previously published there. There is also a 2015 paper by Hernández
//! and López on using AVX2 for the X25519 Montgomery ladder, but
//! neither the paper nor the code are publicly available, and it apparently
//! gives only a [slight speedup][avx2trac].
//! //!
//! HWCD also suggest using a mixed representation, passing between \\( //! HWCD also suggest using a mixed representation, passing between \\(
//! \mathbb P\^3 \\) "extended" coordinates and \\( \mathbb P\^2 \\) //! \mathbb P\^3 \\) "extended" coordinates and \\( \mathbb P\^2 \\)
@ -118,9 +121,12 @@
//! //!
//! This optimization is not used for the parallel formulas, which are //! This optimization is not used for the parallel formulas, which are
//! therefore slightly less efficient when counting the total number of //! therefore slightly less efficient when counting the total number of
//! multiplications and squarings. In addition, the parallel formulas //! field multiplications and squarings. In particular, vectorized doublings
//! can only use a \\( 32 \times 32 \rightarrow 64 \\)-bit multiplier //! are less efficient than serial doublings.
//! instead of a \\( 64 \times 64 \rightarrow 128\\)-bit multiplier. //! In addition, the parallel formulas can only use a \\( 32 \times 32
//! \rightarrow 64 \\)-bit integer multiplier, so the speedup from
//! vectorization must overcome the disadvantage of losing the \\( 64
//! \times 64 \rightarrow 128\\)-bit (serial) integer multiplier.
//! //!
//! When used for constant-time variable-base scalar multiplication, //! When used for constant-time variable-base scalar multiplication,
//! this strategy (using AVX2) gives a significant speedup over the //! this strategy (using AVX2) gives a significant speedup over the
@ -305,6 +311,10 @@
//! The implementation uses the unstable `stdsimd` crate to provide AVX2 //! The implementation uses the unstable `stdsimd` crate to provide AVX2
//! intrinsics, and the code is not yet cleanly factored between the //! intrinsics, and the code is not yet cleanly factored between the
//! field element parts and the point parts. //! field element parts and the point parts.
//!
//! [sandy2x]: https://eprint.iacr.org/2015/943.pdf
//! [avx2trac]: https://trac.torproject.org/projects/tor/ticket/8897#comment:28
//! [hwcd08]: https://www.iacr.org/archive/asiacrypt2008/53500329/53500329.pdf
pub(crate) mod field; pub(crate) mod field;