From f8c5c2e28e5d0457a06fbcaf7b542824523e5722 Mon Sep 17 00:00:00 2001 From: Jack Grigg Date: Mon, 11 Jan 2021 20:34:30 +0000 Subject: [PATCH] book: Add background material Co-authored-by: Sean Bowe Co-authored-by: Daira Hopwood Co-authored-by: ying tong --- book/src/SUMMARY.md | 8 + book/src/background.md | 5 + book/src/background/curves.md | 225 +++++++++++++++++++++++ book/src/background/fields.md | 252 ++++++++++++++++++++++++++ book/src/background/groups.md | 94 ++++++++++ book/src/background/pc-ipa.md | 80 +++++++++ book/src/background/polynomials.md | 276 +++++++++++++++++++++++++++++ book/src/background/recursion.md | 26 +++ book/src/background/upa.md | 77 ++++++++ 9 files changed, 1043 insertions(+) create mode 100644 book/src/background.md create mode 100644 book/src/background/curves.md create mode 100644 book/src/background/fields.md create mode 100644 book/src/background/groups.md create mode 100644 book/src/background/pc-ipa.md create mode 100644 book/src/background/polynomials.md create mode 100644 book/src/background/recursion.md create mode 100644 book/src/background/upa.md diff --git a/book/src/SUMMARY.md b/book/src/SUMMARY.md index 5f11e6f..a720c8a 100644 --- a/book/src/SUMMARY.md +++ b/book/src/SUMMARY.md @@ -19,3 +19,11 @@ - [Gadgets](design/gadgets.md) - [SHA-256](design/gadgets/sha256.md) - [16-bit table chip](design/gadgets/sha256/table16.md) +- [Background Material](background.md) + - [Fields](background/fields.md) + - [Polynomials](background/polynomials.md) + - [Cryptographic groups](background/groups.md) + - [Elliptic curves](background/curves.md) + - [UltraPLONK arithmetisation](background/upa.md) + - [Polynomial commitment using inner product argument](background/pc-ipa.md) + - [Recursion](background/recursion.md) diff --git a/book/src/background.md b/book/src/background.md new file mode 100644 index 0000000..72339d0 --- /dev/null +++ b/book/src/background.md @@ -0,0 +1,5 @@ +# Background Material + +This section covers the background material required to understand the Halo 2 proving +system. It is targeted at an ELI15 (Explain It Like I'm 15) level; if you think anything +could do with additional explanation, let us know! diff --git a/book/src/background/curves.md b/book/src/background/curves.md new file mode 100644 index 0000000..2668740 --- /dev/null +++ b/book/src/background/curves.md @@ -0,0 +1,225 @@ +# Elliptic curves + +Elliptic curves constructed over finite fields are another important cryptographic tool. +There are several ways to define the curve equation, but for our purposes, let +$\mathbb{F}_p$ be a large (255-bit) field, and then let the set of solutions $(x, y)$ to +$y^2 = x^3 + b$ for some constant $b$ define the $\mathbb{F}_p$-rational points on an +elliptic curve $E(\mathbb{F}_p)$. These $(x, y)$ coordinates are called "affine +coordinates". Each of the $\mathbb{F}_p$-rational points, together with a "point at +infinity" $\mathcal{O}$ that serves as the group identity, can be interpreted as an +element of a group. By convention, elliptic curve groups are written additively. + +![](https://i.imgur.com/JvLS6yE.png) +*"Three points on a line sum to zero, which is the point at infinity."* + +The group addition law is simple: to add two points together, find the line that +intersects both points and obtain the third point, and then negate its $y$-coordinate. The +case that a point is being added to itself, called point doubling, requires special +handling: we find the line tangent to the point, and then find the single other point that +intersects this line and then negate. Otherwise, in the event that a point is being +"added" to its negation, the result is the point at infinity. + +The ability to add and double points naturally gives us a way to scale them by integers. +The number of points on the curve is known as the "group order". If this number is prime +$q$, we call the numbers that we scale curve points by "scalars" and consider that they +are all elements of a scalar field $\mathbb{F}_q$. + +Elliptic curves, when properly designed, have an important security property. Given two +random elements $G, H \in E(\mathbb{F}_p)$ finding $a$ such that $[a] G = H$, otherwise +known as the discrete log of $H$ with respect to $G$, is considered computationally +infeasible with classical computers. This is called the elliptic curve discrete log +assumption. + +## Curve arithmetic + +### Point doubling + +The simplest situation is doubling a point $(x_0, y_0)$. Continuing with our example +$y^2 = x^3 + b$, this is done first by computing the derivative +$$ +\lambda = \frac{dy}{dx} = \frac{3x^2}{2y}. +$$ + +To obtain expressions for $(x_1, y_1) = (x_0, y_0) + (x_0, y_0),$ we consider + +$$ +\begin{aligned} +\frac{-y_1 - y_0}{x_1 - x_0} = \lambda &\implies -y_1 = \lambda(x_1 - x_0) + y_0 \\ +&\implies \boxed{y_1 = \lambda(x_0 - x_1) - y_0}. +\end{aligned} +$$ + +To get the expression for $x_1,$ we substitute $y = \lambda(x_0 - x) - y_0$ into the +elliptic curve equation: + +$$ +\begin{aligned} +y^2 = x^3 + b &\implies (\lambda(x_0 - x) - y_0)^2 = x^3 + b \\ +&\implies x^3 - \lambda^2 x^2 + \cdots = 0 \leftarrow\text{(rearranging terms)} \\ +&= (x - x_0)(x - x_0)(x - x_1) \leftarrow\text{(known roots $x_0, x_0, x_1$)} \\ +&= x^3 - (x_0 + x_0 + x_1)x^2 + \cdots. +\end{aligned} +$$ + +Comparing coefficients for the $x^2$ term gives us +$\lambda^2 = x_0 + x_0 + x_1 \implies \boxed{x_1 = \lambda^2 - 2x_0}.$ + + +### Projective coordinates +This unfortunately requires an expensive inversion of $2y$. We can avoid this by arranging +our equations to "defer" the computation of the inverse, since we often do not need the +actual affine $(x', y')$ coordinate of the resulting point immediately after an individual +curve operation. Let's introduce a third coordinate $Z$ and scale our curve equation by +$Z^3$ like so: + +$$ +Z^3 y^2 = Z^3 x^3 + Z^3 b +$$ + +Our original curve is just this curve at the restriction $Z = 1$. If we allow the affine +point $(x, y)$ to be represented by $X = xZ$, $Y = yZ$ and $Z \neq 0$ then we have the +[homogenous projective curve](https://en.wikipedia.org/wiki/Homogeneous_coordinates) + +$$ +Y^2 Z = X^3 + Z^3 b. +$$ + +Obtaining $(x, y)$ from $(X, Y, Z)$ is as simple as computing $(X/Z, Y/Z)$ when +$Z \neq 0$. (When $Z = 0,$ we are dealing with the point at infinity $O := (0:1:0)$.) In +this form, we now have a convenient way to defer the inversion required by doubling a +point. The general strategy is to express $x', y'$ as rational functions using $x = X/Z$ +and $y = Y/Z$, rearrange to make their denominators the same, and then take the resulting +point $(X, Y, Z)$ to have $Z$ be the shared denominator and $X = x'Z, Y = y'Z$. + +> Projective coordinates are often, but not always, more efficient than affine +> coordinates. There may be exceptions to this when either we have a different way to +> apply Montgomery's trick, or when we're in the circuit setting where multiplications and +> inversions are about equally as expensive (at least in terms of circuit size). + +The following shows an example of doubling a point $(X, Y, Z) = (xZ, yZ, Z)$ without an +inversion. Substituting with $X, Y, Z$ gives us +$$ +\lambda = \frac{3x^2}{2y} = \frac{3(X/Z)^2}{2(Y/Z)} = \frac{3 X^2}{2YZ} +$$ + +and gives us +$$ +\begin{aligned} +x' &= \lambda^2 - 2x \\ +&= \lambda^2 - \frac{2X}{Z} \\ +&= \frac{9 X^4}{4Y^2Z^2} - \frac{2X}{Z} \\ +&= \frac{9 X^4 - 8XY^2Z}{4Y^2Z^2} \\ +&= \frac{18 X^4 Y Z - 16XY^3Z^2}{8Y^3Z^3} \\ +\\ +y' &= \lambda (x - x') - y \\ +&= \lambda (\frac{X}{Z} - \frac{9 X^4 - 8XY^2Z}{4Y^2Z^2}) - \frac{Y}{Z} \\ +&= \frac{3 X^2}{2YZ} (\frac{X}{Z} - \frac{9 X^4 - 8XY^2Z}{4Y^2Z^2}) - \frac{Y}{Z} \\ +&= \frac{3 X^3}{2YZ^2} - \frac{27 X^6 - 24X^3Y^2Z}{8Y^3Z^3} - \frac{Y}{Z} \\ +&= \frac{12 X^3Y^2Z - 8Y^4Z^2 - 27 X^6 + 24X^3Y^2Z}{8Y^3Z^3} +\end{aligned} +$$ + +Notice how the denominators of $x'$ and $y'$ are the same. Thus, instead of computing +$(x', y')$ we can compute $(X, Y, Z)$ with $Z = 8Y^3Z^3$ and $X, Y$ set to the +corresponding numerators such that $X/Z = x'$ and $Y/Z = y'$. This completely avoids the +need to perform an inversion when doubling, and something analogous to this can be done +when adding two distinct points. + +### TODO: Point addition +$$ +\begin{aligned} +P + Q &= R\\ +(x_p, y_p) + (x_q, y_q) &= (x_r, y_r) \\ +\lambda &= \frac{y_q - y_p}{x_q - x_p} \\ +x_r &= \lambda^2 - x_p - x_q \\ +y_r &= \lambda(x_p - x_r) - y_p +\end{aligned} +$$ + +---------- + +Important notes: + +* There exist efficient formulae[^complete-formulae] for point addition that do not have + edge cases (so-called "complete" formulae) and that unify the addition and doubling + cases together. The result of adding a point to its negation using those formulae + produces $Z = 0$, which represents the point at infinity. +* In addition, there are other models like the Jacobian representation where + $(x, y) = (xZ^2, yZ^3, Z)$ where the curve is rescaled by $Z^6$ instead of $Z^3$, and + this representation has even more efficient arithmetic but no unified/complete formulae. +* We can easily compare two curve points $(X_1, Y_1, Z_1)$ and $(X_2, Y_2, Z_2)$ for + equality in the homogenous projective coordinate space by "homogenizing" their + Z-coordinates; the checks become $X_1 Z_2 = X_2 Z_1$ and $Y_1 Z_2 = Y_2 Z_1$. + +## Curve endomorphisms + +Imagine that $\mathbb{F}_p$ has a primitive cube root of unity, or in other words that +$3 | p - 1$ and so an element $\zeta_p$ generates a $3$-order multiplicative subgroup. +Notice that a point $(x, y)$ on our example elliptic curve $y^2 = x^3 + b$ has two cousin +points: $(\zeta_p x, \zeta_p^2 x)$, because the computation $x^3$ effectively kills the +$\zeta$ component of the $x$-coordinate. Applying the map $(x, y) \mapsto (\zeta_p x, y)$ +is an application of an endomorphism over the curve. The exact mechanics involved are +complicated, but when the curve has a prime $q$ number of points (and thus a prime +"order") the effect of the endomorphism is to multiply the point by a scalar in +$\mathbb{F}_q$ which is also a primitive cube root $\zeta_q$ in the scalar field. + +## Curve point compression +TODO + +## Cycles of curves +Let $E_p$ be an elliptic curve over a finite field $\mathbb{F}_p,$ where $p$ is a prime. +We denote this by $E_p/\mathbb{F}_p.$ and we denote the group of points of $E_p$ over +$\mathbb{F}_p,$ with order $q = \#E(\mathbb{F}_p).$ For this curve, we call $\mathbb{F}_p$ +the "base field" and $\mathbb{F}_q$ the "scalar field". + +We instantiate our proof system over the elliptic curve $E_p/\mathbb{F}_p$. This allows us +to prove statements about $\mathbb{F}_q$-arithmetic circuit satisfiability. + +> **(aside) If our curve $E_p$ is over $\mathbb{F}_p,$ why is the arithmetic circuit instead in $\mathbb{F}_q$?** +> The proof system is basically working on encodings of the scalars in the circuit (or +> more precisely, commitments to polynomials whose coefficients are scalars). The scalars +> are in $\mathbb{F}_q$ when their encodings/commitments are elliptic curve points in +> $E_p/\mathbb{F}_p$. + +However, most of the verifier's arithmetic computations are over the base field +$\mathbb{F}_p,$ and are thus efficiently expressed as an $\mathbb{F}_p$-arithmetic +circuit. + +> **(aside) Why are the verifier's computations (mainly) over $\mathbb{F}_p$?** +> The Halo 2 verifier actually has to perform group operations using information output by +> the circuit. Group operations like point doubling and addition use arithmetic in +> $\mathbb{F}_p$, because the coordinates of points are in $\mathbb{F}_p.$ + +This motivates us to construct another curve with scalar field $\mathbb{F}_p$, which has +an $\mathbb{F}_p$-arithmetic circuit that can efficiently verify proofs from the first +curve. As a bonus, if this second curve had base field $E_q/\mathbb{F}_q,$ it would +generate proofs that could be efficiently verified in the first curve's +$\mathbb{F}_q$-arithmetic circuit. In other words, we instantiate a second proof system +over $E_q/\mathbb{F}_q,$ forming a 2-cycle with the first: + +![](https://i.imgur.com/bNMyMRu.png) + +### TODO: Pallas-Vesta curves +Reference: https://github.com/zcash/pasta + +## Hashing to curves + +Sometimes it is useful to be able to produce a random point on an elliptic curve +$E_p/\mathbb{F}_p$ corresponding to some input, in such a way that no-one will know its +discrete logarithm (to any other base). + +This is described in detail in the [Internet draft on Hashing to Elliptic Curves][cfrg-hash-to-curve]. +Several algorithms can be used depending on efficiency and security requirements. The +framework used in the Internet Draft makes use of several functions: + +* ``hash_to_field``: takes a byte sequence input and maps it to a element in the base + field $\mathbb{F}_p$ +* ``map_to_curve``: takes an $\mathbb{F}_p$ element and maps it to $E_p$. + +[cfrg-hash-to-curve]: https://datatracker.ietf.org/doc/draft-irtf-cfrg-hash-to-curve/?include_text=1 + +### TODO: Simplified SWU +Reference: https://eprint.iacr.org/2019/403.pdf + +## References +[^complete-formulae]: Renes, J., Costello, C., & Batina, L. (2016, May). "Complete addition formulas for prime order elliptic curves." In Annual International Conference on the Theory and Applications of Cryptographic Techniques (pp. 403-428). Springer, Berlin, Heidelberg. https://eprint.iacr.org/2015/1060.pdf diff --git a/book/src/background/fields.md b/book/src/background/fields.md new file mode 100644 index 0000000..cdcefca --- /dev/null +++ b/book/src/background/fields.md @@ -0,0 +1,252 @@ +# Fields + +A fundamental component of many cryptographic protocols is the algebraic structure known +as a [field]. Fields are sets of objects (usually numbers) with two associated binary +operators $+$ and $\times$ such that various [field axioms][field-axioms] hold. The real +numbers $\mathbb{R}$ are an example of a field with an uncountably infinite number of +elements. + +[field]: https://en.wikipedia.org/wiki/Field_(mathematics) +[field-axioms]: https://en.wikipedia.org/wiki/Field_(mathematics)#Classic_definition + +Halo makes use of _finite fields_ which have a finite number of elements. Finite fields +are fully classified as follows: +- if $\mathbb{F}$ is a finite field, it contains $|\mathbb{F}| = p^k$ elements for some + integer $k \geq 1$ and some prime $p$; +- any two finite fields with the same number of elements are isomorphic. In particular, + all of the arithmetic in a prime field $\mathbb{F}_p$ is isomorphic to addition and + multiplication of integers modulo $p$, i.e. in $\mathbb{Z}_p$. This is why we often + refer to $p$ as the _modulus_. + +We'll write a field as $\mathbb{F}_q$ where $q = p^k$. The prime $p$ is called its +_characteristic_. In the cases where $k \gt 1$ the field $\mathbb{F}_q$ is a $k$-degree +extension of the field $\mathbb{F}_p$. (By analogy, the complex numbers +$\mathbb{C} = \mathbb{R}(i)$ are an extension of the real numbers.) However, in Halo we do +not care about extension fields. Whenever we write $\mathbb{F}_p$ we are referring to what +we call a _prime field_ which has a prime $p$ number of elements, i.e. $k = 1$. + +Important notes: +* There are two special elements in any field: $\mathcal{O}$, the additive identity, and + $1$, the multiplicative identity. +* The least significant bit of a field element, when represented as an integer in binary + format, can be interpreted as its "sign" to help distinguish it from its additive + inverse (negation). This is because for some nonzero element $a$ which has a least + significant bit $0$ we have that $-a = p - a$ has a least significant bit $1$, and vice + versa. We could also use whether or not an element is larger than $(p - 1) / 2$ to give + it a "sign." + +## Inverses and groups + +Any non-zero element $a \in \mathbb{F}_p$ has a _multiplicative inverse_ $b = a^{-1}$, +which is the _unique_ element $b$ such that $ab = 1$. + +A quick way of obtaining the inverse is $a^{-1} = a^{p - 2}$. The reason for this stems +from [Fermat's little theorem][fermat-little], which states that $a^p = a \pmod p$ for any +integer $a$. If $a$ is nonzero, we can divide by $a$ twice to get $a^{p-2} = a^{-1}.$ + +[fermat-little]: https://en.wikipedia.org/wiki/Fermat%27s_little_theorem + +However, it may be more intuitive to understand the set of nonzero elements of +$\mathbb{F}_p$ as a [group], where the group operation is given by multiplication on the +field. We use the notation $\mathbb{F}_p^\times$ for the multiplicative group over the set +$\mathbb{F}_p - \{\mathcal{O}\}$. Groups are simpler and more limited than fields; they +have only _one_ operator $\cdot$ and fewer axioms. + +[group]: https://en.wikipedia.org/wiki/Group_(mathematics) + +> #### (aside) Additive vs multiplicative notation +> If $\cdot$ is written as $+$ and the identity as $0$ or $\mathcal{O}$, then we say the +> group is "written additively". If $\cdot$ is written as $\times$ or omitted (i.e. +> $a \cdot b$ written as $ab$) and the identity as $1$, then we say it is "written +> multiplicatively". It's conventional to use additive notation for +> [elliptic curve groups](curves.md), and multiplicative notation when (as in this case) +> the elements come from a finite field. When additive notation is used, we also write +> +> $$[k] A = \underbrace{A + A + \cdots + A}_{k \text{ times}}$$ +> +> for nonnegative $k$ and call this "scalar multiplication"; we also often use uppercase +> letters for variables denoting group elements. When multiplicative notation is used, we +> also write +> +> $$a^k = \underbrace{a \times a \times \cdots \times a}_{k \text{ times}}$$ +> +> and call this "exponentiation". In either case we call the scalar $k$ such that +> $[k] g = a$ or $g^k = a$ the "discrete logarithm" of $a$ to base $g$. We can extend +> scalars to negative integers by inversion, i.e. $[-k] A + [k] A = \mathcal{O}$ or +> $a^{-k} \times a^k = 1$. + +The _order_ of a group element $a$ is defined (in multiplicative notation) as the smallest +positive integer $k$ such that $a^k = 1$. + +Groups always have [generators] which are elements that, when the group operation is +applied repeatedly with the same element some number of times, produce every other element +of the group. That is, a generator has maximal order, which we also call the order of the +group. + +[generators]: https://en.wikipedia.org/wiki/Generating_set_of_a_group + +There can be many different generators. Let's assume that $\alpha$ is a generator of +$\mathbb{F}_p^\times$, so it has order $p-1$ (equal to the number of elements in +$\mathbb{F}_p^\times$). Therefore, for any element in $a \in \mathbb{F}_p^\times$ there is +a unique integer $i \in \{0..p-2\}$ such that $a = \alpha^i$. + +Notice that $a \times b$ where $a, b \in \mathbb{F}_p^\times$ can really be interpreted as +$\alpha^i \times \alpha^j$ where $a = \alpha^i$ and $b = \alpha^j$. Indeed, it holds that +$\alpha^i \times \alpha^j = \alpha^{i + j}$ for all $0 \leq i, j \lt p - 1$. As a result +the multiplication of nonzero field elements can be interpreted as addition modulo $p - 1$ +with respect to some fixed generator $\alpha$. The addition just happens "in the exponent." + +This is another way to look at where $a^{p - 2}$ comes from for computing inverses in the +field: + +$$p - 2 \equiv -1 \pmod{p - 1},$$ + +so $a^{p - 2} = a^{-1}$. + +### Montgomery's Trick + +Montgomery's trick, named after Peter Montgomery (RIP) is a way to compute many field +inversions at the same time. Imagine we need to compute the inverses of three nonzero +elements $a, b, c \in \mathbb{F}_p$. Instead, we'll compute the products $x = ab$ and +$y = xc = abc$, and compute the inversion + +$$z = y^{p - 2} = \frac{1}{abc}.$$ + +We can now multiply $z$ by $x$ to obtain $\frac{1}{c}$ and multiply $z$ by $c$ to obtain +$\frac{1}{ab}$, which we can then multiply by $a, b$ to obtain their respective inverses. + +This technique generalizes to arbitrary numbers of field elements with just a single +inversion necessary. + +## Multiplicative subgroups + +A _subgroup_ of a group $G$ with operation $\cdot$, is a subset of elements of $G$ that +also form a group under $\cdot$. + +In the previous section we said that $\alpha$ is a generator of the $p - 1$ order +multiplicative group $\mathbb{F}_p^\times$. This is a _composite_ order group, and so by +the Chinese remainder theorem[^chinese-remainder] it has strict subgroups. As an example +let's imagine that $p = 11$, and so $p - 1$ factors into $5 \cdot 2$. Thus, there is a +generator $\beta$ of the $5$-order subgroup and a generator $\gamma$ of the $2$-order +subgroup. All elements in $\mathbb{F}_p^\times$, therefore, can be written uniquely as +$\beta^i \cdot \gamma^j$ for some $i$ (modulo $5$) and some $j$ (modulo $2$). + +If we have $a = \beta^i \cdot \gamma^j$ notice what happens when we compute + +$$ +a^5 = (\beta^i \cdot \gamma^j)^5 + = \beta^{i \cdot 5} \cdot \gamma^{j \cdot 5} + = \beta^0 \cdot \gamma^{j \cdot 5} + = \gamma^{j \cdot 5}; +$$ + +we have effectively "killed" the $5$-order subgroup component, producing a value in the +$2$-order subgroup. + +[Lagrange's theorem (group theory)][lagrange-group] states that the order of any subgroup +$H$ of a finite group $G$ divides the order of $G$. Therefore, the order of any subgroup +of $\mathbb{F}_p^\times$ must divide $p-1.$ + +[lagrange-group]: https://en.wikipedia.org/wiki/Lagrange%27s_theorem_(group_theory) + +## Square roots + +In a field $\mathbb{F}_p$ exactly half of all nonzero elements are squares; the remainder +are non-squares or "quadratic non-residues". In order to see why, consider an $\alpha$ +that generates the $2$-order multiplicative subgroup of $\mathbb{F}_p^\times$ (this always +exists because $p - 1$ is divisible by $2$ since $p$ is prime) and $\beta$ that generates +the $t$-order multiplicative subgroup of $\mathbb{F}_p^\times$ where $p - 1 = 2t$. Then +every element $a \in \mathbb{F}_p^\times$ can be written uniquely as +$\alpha^i \cdot \beta^j$ with $i \in \mathbb{Z}_2$ and $j \in \mathbb{Z}_t$. Half of all +elements will have $i = 0$ and the other half will have $i = 1$. + +Let's consider the simple case where $p \equiv 3 \pmod{4}$ and so $t$ is odd (if $t$ is +even, then $p - 1$ would be divisible by $4$, which contradicts $p$ being $3 \pmod{4}$). +If $a \in \mathbb{F}_p^\times$ is a square, then there must exist +$b = \alpha^i \cdot \beta^j$ such that $b^2 = a$. But this means that + +$$a = (\alpha^i \cdot \beta^j)^2 = \alpha^{2i} \cdot \beta^{2j} = \beta^{2j}.$$ + +In other words, all squares in this particular field do not generate the $2$-order +multiplicative subgroup, and so since half of the elements generate the $2$-order subgroup +then at most half of the elements are square. In fact exactly half of the elements are +square (since squaring each nonsquare element gives a unique square). This means we can +assume all squares can be written as $\beta^m$ for some $m$, and therefore finding the +square root is a matter of exponentiating by $2^{-1} \pmod{t}$. + +In the event that $p \equiv 1 \pmod{4}$ then things get more complicated because +$2^{-1} \pmod{t}$ does not exist. Let's write $p - 1$ as $2^k \cdot t$ with $t$ odd. The +case $k = 0$ is impossible, and the case $k = 1$ is what we already described, so consider +$k \geq 2$. $\alpha$ generates a $2^k$-order multiplicative subgroup and $\beta$ generates +the odd $t$-order multiplicative subgroup. Then every element $a \in \mathbb{F}_p^\times$ +can be written as $\alpha^i \cdot \beta^j$ for $i \in \mathbb{Z}_{2^k}$ and +$j \in \mathbb{Z}_t$. If the element is a square, then there exists some $b = \sqrt{a}$ +which can be written $b = \alpha^{i'} \cdot \beta^{j'}$ for $i' \in \mathbb{Z}_{2^k}$ and +$j' \in \mathbb{Z}_t$. This means that $a = b^2 = \alpha^{2i'} \cdot \beta^{2j'}$, +therefore we have $i \equiv 2i' \pmod{2^k}$, and $j \equiv 2j' \pmod{t}$. $i$ would have +to be even in this case because otherwise it would be impossible to have +$i \equiv 2i' \pmod{2^k}$ for any $i'$. In the case that $a$ is not a square, then $i$ is +odd, and so half of all elements are squares. + +In order to compute the square root, we can first raise the element +$a = \alpha^i \cdot \beta^j$ to the power $t$ to "kill" the $t$-order component, giving + +$$a^t = \alpha^{it \pmod 2^k} \cdot \beta^{jt \pmod t} = \alpha^{it \pmod 2^k}$$ + +and then raise this result to the power $t^{-1} \pmod{2^k}$ to undo the effect of the +original exponentiation on the $2^k$-order component: + +$$(\alpha^{it \bmod 2^k})^{t^{-1} \pmod{2^k}} = \alpha^i$$ + +(since $t$ is relatively prime to $2^k$). This leaves bare the $\alpha^i$ value which we +can trivially handle. We can similarly kill the $2^k$-order component to obtain +$\beta^{j \cdot 2^{-1} \pmod{t}}$, and put the values together to obtain the square root. + +It turns out that in the cases $k = 2, 3$ there are simpler algorithms that merge several +of these exponentiations together for efficiency. For other values of $k$, the only known +way is to manually extract $i$ by squaring until you obtain the identity for every single +bit of $i$. This is the essence of the [Tonelli-Shanks square root algorithm][ts-sqrt] and +describes the general strategy. (There is another square root algorithm that uses +quadratic extension fields, but it doesn't pay off in efficiency until the prime becomes +quite large.) + +> TODO: describe more recent algorithms, e.g. Bernstein's table-based method and +> [eprint 2020/1407](https://eprint.iacr.org/2020/1407). + +[ts-sqrt]: https://en.wikipedia.org/wiki/Tonelli%E2%80%93Shanks_algorithm + +## Roots of unity + +In the previous sections we wrote $p - 1 = 2^k \cdot t$ with $t$ odd, and stated that an +element $\alpha \in \mathbb{F}_p^\times$ generated the $2^k$-order subgroup. For +convenience, let's denote $n := 2^k.$ The elements $\{1, \alpha, \alpha^2, \alpha^{n-1}\}$ +are known as the $n$th [roots of unity](https://en.wikipedia.org/wiki/Root_of_unity). + +The **primitive root of unity**, $\omega,$ is an $n$th root of unity such that +$\omega^i \neq 1$ except when $i \equiv 0 \pmod{n}$. + +Important notes: +- if $\alpha$ is an $n$th root of unity, $\alpha$ satisfies $\alpha^n - 1 = 0.$ If + $\alpha \neq 1,$ then + $$1 + \alpha + \alpha^2 + \cdots + \alpha^{n-1} = 0.$$ +- equivalently, the roots of unity are solutions to the equation + $$X^n - 1 = (X - 1)(X - \alpha)(X - \alpha^2) \cdots (X - \alpha^{n-1}).$$ +- **$\boxed{\omega^{\frac{n}{2}+i} = -\omega^i}$ ("Negation lemma")**. Proof: + $$ + \begin{aligned} + \omega^n = 1 &\implies \omega^n - 1 = 0 \\ + &\implies (\omega^{n/2} + 1)(\omega^{n/2} - 1) = 0. + \end{aligned} + $$ + Since the order of $\omega$ is $n$, $\omega^{n/2} \neq 1.$ Therefore, $\omega^{n/2} = -1.$ + +- **$\boxed{(\omega^{\frac{n}{2}+i})^2 = (\omega^i)^2}$ ("Halving lemma")**. Proof: + $$ + (\omega^{\frac{n}{2}+i})^2 = \omega^{n + 2i} = \omega^{n} \cdot \omega^{2i} = \omega^{2i} = (\omega^i)^2. + $$ + In other words, if we square each element in the $n$th roots of unity, we would get back + only half the elements, $\{(\omega_n^i)^2\} = \{\omega_{n/2}\}$ (i.e. the $n/2$th roots + of unity). There is a two-to-one mapping between the elements and their squares. + +## References +[^chinese-remainder]: Friedman, R. (n.d.) "Cyclic Groups and Elementary Number Theory II" (p. 5). http://www.math.columbia.edu/~rf/numbertheory2.pdf diff --git a/book/src/background/groups.md b/book/src/background/groups.md new file mode 100644 index 0000000..12b622b --- /dev/null +++ b/book/src/background/groups.md @@ -0,0 +1,94 @@ +# Cryptographic groups + +In the section [Inverses and groups](fields.md#inverses-and-groups) we introduced the +concept of *groups*. A group has an identity and a group operation. In this section we +will write groups additively, i.e. the identity is $\mathcal{O}$ and the group operation +is $+$. + +Some groups can be used as *cryptographic groups*. At the risk of oversimplifying, this +means that the problem of finding a discrete logarithm of a group element $P$ to a given +base $G$, i.e. finding $x$ such that $P = [x] G$, is hard in general. + +## Pedersen commitment +The Pedersen commitment [[P99]] is a way to commit to a secret message in a verifiable +way. It uses two random public generators $G, H \in \mathbb{G},$ where $\mathbb{G}$ is a +cryptographic group of order $p$. A random secret $r$ is chosen in $\mathbb{Z}_q$, and the +message to commit to $m$ is from any subset of $\mathbb{Z}_q$. The commitment is + +$$c = \text{Commit}(m,r)=[m]G + [r]H.$$ + +To open the commitment, the committer reveals $m$ and $r,$ thus allowing anyone to verify +that $c$ is indeed a commitment to $m.$ + +[P99]: https://link.springer.com/content/pdf/10.1007%2F3-540-46766-1_9.pdf#page=3 + +Notice that the Pedersen commitment scheme is homomorphic: + +$$ +\begin{aligned} +\text{Commit}(m,r) + \text{Commit}(m',r') &= [m]G + [r]H + [m']G + [r']H \\ +&= [m + m']G + [r + r']H \\ +&= \text{Commit}(m + m',r + r'). +\end{aligned} +$$ + +Assuming the discrete log assumption holds, Pedersen commitments are also perfectly hiding +and computationally binding: + +* **hiding**: the adversary chooses messages $m_0, m_1.$ The committer commits to one of + these messages $c = \text{Commit}(m_b;r), b \in \{0,1\}.$ Given $c,$ the probability of + the adversary guessing the correct $b$ is no more than $\frac{1}{2}$. +* **binding**: the adversary cannot pick two different messages $m_0 \neq m_1,$ and + randomness $r_0, r_1,$ such that $\text{Commit}(m_0,r_0) = \text{Commit}(m_1,r_1).$ + +### Vector Pedersen commitment +We can use a variant of the Pedersen commitment scheme to commit to multiple messages at +once, $\mathbf{m} = (m_1, \cdots, m_n)$. This time, we'll have to sample a corresponding +number of random public generators $\mathbf{G} = (G_0, \cdots, G_{n-1}),$ along with a +single random generator $H$ as before (for use in hiding). Then, our commitment scheme is: + +$$ +\begin{aligned} +\text{Commit}(\mathbf{m}; r) &= \text{Commit}((m_0, \cdots, m_{n-1}); r) \\ +&= [r]H + [m_0]G_0 + \cdots + [m_{n-1}]G_{n-1} \\ +&= [r]H + \sum_{i= 0}^{n-1} [m_i]G_i. +\end{aligned} +$$ + +> TODO: is this positionally binding? + +## Diffie--Hellman + +An example of a protocol that uses cryptographic groups is Diffie--Hellman key agreement +[[DH1976]]. The Diffie--Hellman protocol is a method for two users, Alice and Bob, to +generate a shared private key. It proceeds as follows: + +1. Alice and Bob publicly agree on two prime numbers, $p$ and $G,$ where $p$ is large and + $G$ is a primitive root $\pmod p.$ (Note that $g$ is a generator of the group + $\mathbb{F}_p^\times.$) +2. Alice chooses a large random number $a$ as her private key. She computes her public key + $A = [a]G \pmod p,$ and sends $A$ to Bob. +3. Similarly, Bob chooses a large random number $b$ as his private key. He computes his + public key $B = [b]G \pmod p,$ and sends $B$ to Alice. +4. Now both Alice and Bob compute their shared key $K = [ab]G \pmod p,$ which Alice + computes as + $$K = [a]B \pmod p = [a]([b]G) \pmod p,$$ + and Bob computes as + $$K = [b]A \pmod p = [b]([a]G) \pmod p.$$ + +[DH1976]: https://ee.stanford.edu/~hellman/publications/24.pdf + +A potential eavesdropper would need to derive $K = [ab]g \pmod p$ knowing only +$g, p, A = [a]G,$ and $B = [b]G$: in other words, they would need to either get the +discrete logarithm $a$ from $A = [a]G$ or $b$ from $B = [b]G,$ which we assume to be +computationally infeasible in $\mathbb{F}_p^\times.$ + +More generally, protocols that use similar ideas to Diffie--Hellman are used throughout +cryptography. One way of instantiating a cryptographic group is as an +[elliptic curve](curves.md). Before we go into detail on elliptic curves, we'll describe +some algorithms that can be used for any group. + +## Multiscalar multiplication + +### TODO: Pippenger's algorithm +Reference: https://jbootle.github.io/Misc/pippenger.pdf diff --git a/book/src/background/pc-ipa.md b/book/src/background/pc-ipa.md new file mode 100644 index 0000000..5388856 --- /dev/null +++ b/book/src/background/pc-ipa.md @@ -0,0 +1,80 @@ +# Polynomial commitment using inner product argument +We want to commit to some polynomial $p(X) \in \mathbb{F}_p[X]$, and be able to provably +evaluate the committed polynomial at arbitrary points. The naive solution would be for the +prover to simply send the polynomial's coefficients to the verifier: however, this +requires $O(n)$ communication. Our polynomial commitment scheme gets the job done using +$O(\log n)$ communication. + +### `Setup` +Given a parameter $d = 2^k,$ we generate the common reference string +$\sigma = (\mathbb{G}, \mathbf{G}, H, \mathbb{F}_p)$ defining certain constants for this +scheme: +* $\mathbb{G}$ is a group of prime order $p;$ +* $\mathbf{G} \in \mathbb{G}^d$ is a vector of $d$ random group elements; +* $H \in \mathbb{G}$ is a random group element; and +* $\mathbb{F}_p$ is the finite field of order $p.$ + +### `Commit` +The Pedersen vector commitment $\text{Commit}$ is defined as + +$$\text{Commit}(\sigma, p(X); r) = \langle\mathbf{a}, \mathbf{G}\rangle + [r]H,$$ + +for some polynomial $p(X) \in \mathbb{F}_p[X]$ and some blinding factor +$r \in \mathbb{F}_p.$ Here, each element of the vector $\mathbf{a}_i \in \mathbb{F}_p$ is +the coefficient for the $i$th degree term of $p(X),$ and $p(X)$ is of maximal degree +$d - 1.$ + +### `Open` (prover) and `OpenVerify` (verifier) +The modified inner product argument is an argument of knowledge for the relation + +$$\boxed{\{((P, x, v); (\mathbf{a}, r)): P = \langle\mathbf{a}, \mathbf{G}\rangle + [r]H, v = \langle\mathbf{a}, \mathbf{b}\rangle\}},$$ + +where $\mathbf{b} = (1, x, x^2, \cdots, x^{d-1})$ is composed of increasing powers of the +evaluation point $x.$ This allows a prover to demonstrate to a verifier that the +polynomial contained “inside” the commitment $P$ evaluates to $v$ at $x,$ and moreover, +that the committed polynomial has maximum degree $d − 1.$ + +The inner product argument proceeds in $k = \log_2 d$ rounds. For our purposes, it is +sufficient to know about its final outputs, while merely providing intuition about the +intermediate rounds. (Refer to Section 3 in the [Halo] paper for a full explanation.) + +[Halo]: https://eprint.iacr.org/2019/1021.pdf + +Before beginning the argment, the verifier selects a random group element $U$ and sends it +to the prover. We initialise the argument at round $k,$ with the vectors +$\mathbf{a}^{(k)} := \mathbf{a},$ $\mathbf{G}^{(k)} := \mathbf{G}$ and +$\mathbf{b}^{(k)} := \mathbf{b}.$ In each round $j = k, k-1, \cdots, 1$: + +* the prover computes two values $L_j$ and $R_j$ by taking some inner product of + $\mathbf{a}^{(j)}$ with $\mathbf{G}^{(j)}$ and $\mathbf{b}^{(j)}$. Note that are in some + sense "cross-terms": the lower half of $\mathbf{a}$ is used with the higher half of + $\mathbf{G}$ and $\mathbf{b}$, and vice versa: + +$$ +\begin{aligned} +L_j &= \langle\mathbf{a_{lo}^{(j)}}, \mathbf{G_{hi}^{(j)}}\rangle + [l_j]H + [\langle\mathbf{a_{lo}^{(j)}}, \mathbf{b_{hi}^{(j)}}\rangle] U\\ +R_j &= \langle\mathbf{a_{hi}^{(j)}}, \mathbf{G_{lo}^{(j)}}\rangle + [l_j]H + [\langle\mathbf{a_{hi}^{(j)}}, \mathbf{b_{lo}^{(j)}}\rangle] U\\ +\end{aligned} +$$ + +* the verifier issues a random challenge $u_j$; +* the prover uses $u_j$ to compress the lower and higher halves of $\mathbf{a}^{(j)}$, + thus producing a new vector of half the original length + $$\mathbf{a}^{(j-1)} = \mathbf{a_{hi}^{(j)}}\cdot u_j^{-1} + \mathbf{a_{lo}^{(j)}}\cdot u_j.$$ + The vectors $\mathbf{G}^{(j)}$ and $\mathbf{b}^{(j)}$ are similarly compressed to give + $\mathbf{G}^{(j-1)}$ and $\mathbf{b}^{(j-1)}$. +* $\mathbf{a}^{(j-1)}$, $\mathbf{G}^{(j-1)}$ and $\mathbf{b}^{(j-1)}$ are input to the + next round $j - 1.$ + +Note that at the end of the last round $j = 1,$ we are left with $a := \mathbf{a}^{(0)}$, +$G := \mathbf{G}^{(0)}$, $b := \mathbf{b}^{(0)},$ each of length 1. The intuition is that +these final scalars, together with the challenges $\{u_j\}$ and "cross-terms" +$\{L_j, R_j\}$ from each round, encode the compression in each round. Since the prover did +not know the challenges $U, \{u_j\}$ in advance, they would have been unable to manipulate +the round compressions. Thus, checking a constraint on these final terms should enforce +that the compression had been performed correctly, and that the original $\mathbf{a}$ +satisfied the relation before undergoing compression. + +Note that $G, b$ are simply rearrangements of the publicly known $\mathbf{G}, \mathbf{b},$ +with the round challenges $\{u_j\}$ mixed in: this means the verifier can compute $G, b$ +independently and verify that the prover had provided those same values. diff --git a/book/src/background/polynomials.md b/book/src/background/polynomials.md new file mode 100644 index 0000000..230fba6 --- /dev/null +++ b/book/src/background/polynomials.md @@ -0,0 +1,276 @@ +# Polynomials + +Let $A(X)$ be a polynomial over $\mathbb{F}_p$ with formal indeterminate $X$. As an example, + +$$ +A(X) = a_0 + a_1 X + a_2 X^2 + a_3 X^3 +$$ + +defines a degree-$3$ polynomial. $a_0$ is referred to as the constant term. Polynomials of +degree $n-1$ have $n$ coefficients. + +Important notes: + +* Multiplication of polynomials produces a product polynomial that is the sum of the + degrees of its factors. Polynomial division subtracts from the degree. + $$\deg(A(X)B(X)) = \deg(A(X)) + \deg(B(X)),$$ + $$\deg(A(X)/B(X)) = \deg(A(X)) -\deg(B(X)).$$ +* Given a polynomial $A(X)$ of degree $n-1$, if we obtain $n$ evaluations of the + polynomial at distinct points then these evaluations perfectly define the polynomial. In + other words, given these evaluations we can obtain a unique polynomial $A(X)$ of degree + $n-1$ via polynomial interpolation. +* $[a_0, a_1, \cdots, a_{n-1}]$ is the **coefficient representation** of the polynomial + $A(X)$. Equivalently, we could use its **evaluation representation** + $$[(x_0, A(x_0)), (x_1, A(x_1)), \cdots, (x_{n-1}, A(x_{n-1}))]$$ + at $n$ distinct points. Either representation uniquely specifies the same polynomial. + +> #### (aside) Horner's rule +> Horner's rule allows for efficient evaluation of a polynomial of degree $n-1$, using +> only $n-1$ multiplications and $n-1$ additions. It is the following identity: +> $$\begin{aligned}a_0 &+ a_1X + a_2X^2 + \cdots + a_{n-1}X^{n-1} \\ &= a_0 + X\bigg( a_1 + X \Big( a_2 + \cdots + X(a_{n-2} + X a_{n-1}) \Big)\!\bigg),\end{aligned}$$ + +## Fast Fourier Transform (FFT) +The FFT is an efficient way of converting between the coefficient and evaluation +representations of a polynomial. It evaluates the polynomial at the $n$th roots of unity +$\{\omega^0, \omega^1, \cdots, \omega^{n-1}\},$ where $\omega$ is a primitive $n$th root +of unity. By exploiting symmetries in the roots of unity, each round of the FFT reduces +the evaluation into a problem only half the size. Most commonly we use polynomials of +length some power of two, $n = 2^k$, and apply the halving reduction recursively. + +### Motivation: Fast polynomial multiplication +In the coefficient representation, it takes $O(n^2)$ operations to multiply two +polynomials $A(X)\cdot B(X) = C(X)$: + +$$ +\begin{aligned} +A(X) &= a_0 + a_1X + a_2X^2 + \cdots + a_{n-1}X^{n-1}, \\ +B(X) &= b_0 + b_1X + b_2X^2 + \cdots + b_{n-1}X^{n-1}, \\ +C(X) &= a_0\cdot (b_0 + b_1X + b_2X^2 + \cdots + b_{n-1}X^{n-1}) \\ +&+ a_1X\cdot (b_0 + b_1X + b_2X^2 + \cdots + b_{n-1}X^{n-1})\\ +&+ \cdots \\ +&+ a_{n-1}X^{n-1} \cdot (b_0 + b_1X + b_2X^2 + \cdots + b_{n-1}X^{n-1}), +\end{aligned} +$$ + +where each of the $n$ terms in the first polynomial has to be multiplied by the $n$ terms +of the second polynomial. + +In the evaluation representation, however, polynomial multiplication only requires $O(n)$ +operations: + +$$ +\begin{aligned} +A&: \{(x_0, A(x_0)), (x_1, A(x_1)), \cdots, (x_{n-1}, A(x_{n-1}))\}, \\ +B&: \{(x_0, B(x_0)), (x_1, B(x_1)), \cdots, (x_{n-1}, B(x_{n-1}))\}, \\ +C&: \{(x_0, A(x_0)B(x_0)), (x_1, A(x_1)B(x_1)), \cdots, (x_{n-1}, A(x_{n-1})B(x_{n-1}))\}, +\end{aligned} +$$ + +where each evaluation is multiplied pointwise. + +This suggests the following strategy for fast polynomial multiplication: + +1. Evaluate polynomials at all $n$ points; +2. Perform fast pointwise multiplication in the evaluation representation ($O(n)$); +3. Convert back to the coefficient representation. + +The challenge now is how to **evaluate** and **interpolate** the polynomials efficiently. +Naively, evaluating a polynomial at $n$ points would require $O(n^2)$ operations (we use +the $O(n)$ Horner's rule at each point): + +$$ +\begin{bmatrix} +A(1) \\ +A(\omega) \\ +A(\omega^2) \\ +\vdots \\ +A(\omega^{n-1}) +\end{bmatrix} = +\begin{bmatrix} +1&1&1&\dots&1 \\ +1&\omega&\omega^2&\dots&\omega^{n-1} \\ +1&\omega^2&\omega^{2\cdot2}&\dots&\omega^{2\cdot(n-1)} \\ +\vdots&\vdots&\vdots& &\vdots \\ +1&\omega^{n-1}&\omega^{2(n-1)}&\cdots&\omega^{(n-1)^2}\\ +\end{bmatrix} \cdot +\begin{bmatrix} +a_0 \\ +a_1 \\ +a_2 \\ +\vdots \\ +a_{n-1} +\end{bmatrix}. +$$ + +For convenience, we will denote the matrices above as: +$$\hat{\mathbf{A}} = \mathbf{V}_\omega \cdot \mathbf{A}. $$ + +($\hat{\mathbf{A}}$ is known as the *Discrete Fourier Transform* of $\mathbf{A}$; +$\mathbf{V}_\omega$ is also called the *Vandermonde matrix*.) + +### The (radix-2) Cooley-Tukey algorithm +Our strategy is to divide a DFT of size $n$ into two interleaved DFTs of size $n/2$. Given +the polynomial $A(X) = a_0 + a_1X + a_2X^2 + \cdots + a_{n-1}X^{n-1},$ we split it up into +even and odd terms: + +$$ +\begin{aligned} +A_{\text{even}} &= a_0 + a_2X + \cdots + a_{n-2}X^{\frac{n}{2} - 1}, \\ +A_{\text{odd}} &= a_1 + a_3X + \cdots + a_{n-1}X^{\frac{n}{2} - 1}. \\ +\end{aligned} +$$ + +To recover the original polynomial, we do +$A(X) = A_{\text{even}} (X^2) + X A_{\text{odd}}(X^2).$ + +Trying this out on points $\omega_n^i$ and $\omega_n^{\frac{n}{2} + i}$, +$i \in [0..\frac{n}{2}-1],$ we start to notice some symmetries: + +$$ +\begin{aligned} +A(\omega_n^i) &= A_{\text{even}} ((\omega_n^i)^2) + \omega_n^i A_{\text{odd}}((\omega_n^i)^2), \\ +A(\omega_n^{\frac{n}{2} + i}) &= A_{\text{even}} ((\omega_n^{\frac{n}{2} + i})^2) + \omega_n^{\frac{n}{2} + i} A_{\text{odd}}((\omega_n^{\frac{n}{2} + i})^2) \\ +&= A_{\text{even}} ((-\omega_n^i)^2) - \omega_n^i A_{\text{odd}}((-\omega_n^i)^2) \leftarrow\text{(negation lemma)} \\ +&= A_{\text{even}} ((\omega_n^i)^2) - \omega_n^i A_{\text{odd}}((\omega_n^i)^2). +\end{aligned} +$$ + +Notice that we are only evaluating $A_{\text{even}}(X)$ and $A_{\text{odd}}(X)$ over half +the domain $\{(\omega_n^0)^2, (\omega_n)^2, \cdots, (\omega_n^{\frac{n}{2} -1})^2\} = \{\omega_{n/2}^i\}, i = [0..\frac{n}{2}-1]$ (halving lemma). +This gives us all the terms we need to reconstruct $A(X)$ over the full domain +$\{\omega^0, \omega, \cdots, \omega^{n -1}\}$: which means we have transformed a +length-$n$ DFT into two length-$\frac{n}{2}$ DFTs. + +We choose $n = 2^k$ to be a power of two (by zero-padding if needed), and apply this +divide-and-conquer strategy recursively. By the Master Theorem[^master-thm], this gives us +an evaluation algorithm with $O(n\log_2n)$ operations, also known as the Fast Fourier +Transform (FFT). + +### Inverse FFT +So we've evaluated our polynomials and multiplied them pointwise. What remains is to +convert the product from the evaluation representation back to coefficient representation. +To do this, we simply call the FFT on the evaluation representation. However, this time we +also: +- replace $\omega^i$ by $\omega^{-i}$ in the Vandermonde matrix, and +- multiply our final result by a factor of $1/n$. + +In other words: +$$\mathbf{A} = \frac{1}{n} \mathbf{V}_{\omega^{-1}} \cdot \hat{\mathbf{A}}. $$ + +(To understand why the inverse FFT has a similar form to the FFT, refer to Slide 13-1 of +[^ifft]. The below image was also taken from [^ifft].) + +![](https://i.imgur.com/lSw30zo.png) + + +## The Schwartz-Zippel lemma +The Schwartz-Zippel lemma informally states that "different polynomials are different at +most points." Formally, it can be written as follows: + +> Let $p(x_1, x_2, \cdots, x_n)$ be a nonzero polynomial of $n$ variables with degree $d$. +> Let $S$ be a finite set of numbers with at least $d$ elements in it. If we choose random +> $\alpha_1, \alpha_1, \cdots, \alpha_n$ from $S$, +> $$\text{Pr}[p(\alpha_1, \alpha_2, \cdots, \alpha_n) = 0] \leq \frac{d}{|S|}.$$ + +In the familiar univariate case $p(X)$, this reduces to saying that a nonzero polynomial +of degree $d$ has at most $d$ roots. + +The Schwartz-Zippel lemma is used in polynomial equality testing. Given two multi-variate +polynomials $p_1(x_1,\cdots,x_n)$ and $p_2(x_1,\cdots,x_n)$ of degrees $d_1, d_2$ +respectively, we can test if +$p_1(\alpha_1, \cdots, \alpha_n) - p_2(\alpha_1, \cdots, \alpha_n) = 0$ for random +$\alpha_1, \cdots, \alpha_n \leftarrow S,$ where the size of $S$ is at least +$|S| \geq (d_1 + d_2).$ If the two polynomials are identical, this will always be true, +whereas if the two polynomials are different then the equality holds with probability at +most $\frac{\max(d_1,d_2)}{|S|}$. + +## Vanishing polynomial +Consider the order-$n$ multiplicative subgroup $\mathcal{H}$ with primitive root of unity +$\omega$. For all $\omega^i \in \mathcal{H}, i \in [n-1],$ we have +$(\omega^i)^n = (\omega^n)^i = (\omega^0)^i = 1.$ In other words, every element of +$\mathcal{H}$ fulfils the equation + +$$ +\begin{aligned} +Z_H(X) &= X^n - 1 \\ +&= (X-\omega^0)(X-\omega^1)(X-\omega^2)\cdots(X-\omega^{n-1}), +\end{aligned} +$$ + +meaning every element is a root of $Z_H(X).$ We call $Z_H(X)$ the **vanishing polynomial** +over $\mathcal{H}$ because it evaluates to zero on all elements of $\mathcal{H}.$ + +This comes in particularly handy when checking polynomial constraints. For instance, to +check that $A(X) + B(X) = C(X)$ over $\mathcal{H},$ we simply have to check that +$A(X) + B(X) - C(X)$ is some multiple of $Z_H(X)$. In other words, if dividing our +constraint by the vanishing polynomial still yields some polynomial +$\frac{A(X) + B(X) - C(X)}{Z_H(X)} = H(X),$ we are satisfied that $A(X) + B(X) - C(X) = 0$ +over $\mathcal{H}.$ + +## Lagrange basis functions +Polynomials are commonly written in the monomial basis (e.g. $X, X^2, ... X^n$). However, +when working over a multiplicative subgroup, we find a more natural expression in the +Lagrange basis. + +Consider the order-$n$ multiplicative subgroup $\mathcal{H}$ with primitive root of unity +$\omega$. The Lagrange basis corresponding to this subgroup is a set of functions +$\{\mathcal{L}_i\}_{i = 0}^{n-1}$, where + +$$ +\mathcal{L_i}(\omega^j) = \begin{cases} +1 & \text{if } i = j, \\ +0 & \text{otherwise.} +\end{cases} +$$ + +We can write this more compactly as $\mathcal{L_i}(\omega^j) = \delta_{ij},$ where +$\delta$ is the Kronecker delta function. + +Now, we can write our polynomial as a linear combination of Lagrange basis functions, + +$$A(X) = \sum_{i = 0}^{n-1} a_i\mathcal{L_i}(X), X \in \mathcal{H},$$ + +which is equivalent to saying that $p(X)$ evaluates to $a_0$ at $\omega^0$, +$p(\omega^1) = a_1, p(\omega^2) = a_2, \cdots,$ and so on. + +When working over a multiplicative subgroup, the Lagrange basis function has a convenient +sparse representation of the form + +$$ +\mathcal{L}_i(X) = \frac{c_i\cdot(X^{n} - 1)}{X - \omega^i}, +$$ + +where $c_i$ is the barycentric weight. (To understand how this form was derived, refer to +[^barycentric].) For $i = 0,$ we have +$c = 1/n \implies \mathcal{L}_0(X) = \frac{1}{n} \frac{(X^{n} - 1)}{X - 1}$. + +Since we cannot assume that the $x_i$'s form a multiplicative subgroup, we consider also +the Lagrange polynomials $\mathcal{L}_i$'s in the general case. Given a set of evaluation +points $\{x_0, x_1, \cdots, x_{n-1}\},$ we can construct + +$$ +\mathcal{L}_i(X) = \prod_{j\neq i}\frac{X - x_j}{x_i - x_j}, i \in [0..n-1]. +$$ + +Here, every $X = x_j \neq x_i$ will produce a zero numerator term $(x_j - x_j),$ causing +the whole product to evaluate to zero. On the other hand, $X= x_i$ will evaluate to +$\frac{x_i - x_j}{x_i - x_j}$ at every term, resulting in an overall product of one. This +gives the desired Kronecker delta behaviour $\mathcal{L_i}(x_j) = \delta_{ij}.$ + +### Lagrange interpolation +Given a polynomial in its evaluation representation + +$$A: \{(x_0, A(x_0)), (x_1, A(x_1)), \cdots, (x_{n-1}, A(x_{n-1}))\},$$ + +we can reconstruct its coefficient form in the Lagrange basis: + +$$A(X) = \sum_{i = 0}^{n-1} A(x_i)\mathcal{L_i}(X), $$ + +where $X \in \{x_0, x_1,\cdots, x_{1-n}\}.$ + +## References +[^master-thm]: Dasgupta, S., Papadimitriou, C. H., & Vazirani, U. V. (2008). "Algorithms" (ch. 2). New York: McGraw-Hill Higher Education. https://people.eecs.berkeley.edu/~vazirani/algorithms/chap2.pdf + +[^ifft]: http://www.cs.ust.hk/mjg_lib/Classes/COMP3711H_Fall16/lectures/FFT_Slides.pdf + +[^barycentric]: Berrut, J. and Trefethen, L. (2004). "Barycentric Lagrange Interpolation." https://people.maths.ox.ac.uk/trefethen/barycentric.pdf diff --git a/book/src/background/recursion.md b/book/src/background/recursion.md new file mode 100644 index 0000000..30bf19f --- /dev/null +++ b/book/src/background/recursion.md @@ -0,0 +1,26 @@ +## Recursion +> Alternative terms: Induction; Accumulation scheme; Proof-carrying data + +However, the computation of $G$ requires a length-$2^k$ multiexponentiation +$\langle \mathbf{G}, \mathbf{s}\rangle,$ where $\mathbf{s}$ is composed of the round +challenges $u_1, \cdots, u_k$ arranged in a binary counting structure. This is the +linear-time computation that we want to amortise across a batch of proof instances. +Instead of computing $G,$ notice that we can express $G$ as a commitment to a polynomial + +$$G = \text{Commit}(\sigma, g(X, u_1, \cdots, u_k)),$$ + +where $g(X, u_1, \cdots, u_k) := \prod_{i=1}^k (u_i + u_i^{-1}X^{2^{i-1}})$ is a +polynomial with degree $2^k - 1.$ + +| | | +| -------- | -------- | +| | Since $G$ is a commitment, it can be checked in an inner product argument. The verifier circuit witnesses $G$ and brings $G, u_1, \cdots, u_k$ out as public inputs to the proof $\pi.$ The next verifier instance checks $\pi$ using the inner product argument; this includes checking that $G = \text{Commit}(g(X, u_1, \cdots, u_k))$ evaluates at some random point to the expected value for the given challenges $u_1, \cdots, u_k.$ Recall from the [previous section](#Polynomial-commitment-using-inner-product-argument) that this check only requires $\log d$ work.

At the end of checking $\pi$ and $G,$ the circuit is left with a new $G',$ along with the $u_1', \cdots, u_k'$ challenges sampled for the check. To fully accept $\pi$ as valid, we should perform a linear-time computation of $G' = \langle\mathbf{G}, \mathbf{s}'\rangle$. Once again, we delay this computation by witnessing $G'$ and bringing $G, u_1, \cdots, u_k$ out as public inputs to the proof $\pi.$

This goes on from one proof instance to the next, until we are satisfied with the size of our batch of proofs. We finally perform a single linear-time computation, thus deciding the validity of the whole batch. | + +We recall from the section [Cycles of curves](curves.md#cycles-of-curves) that we can +instantiate this protocol over a two-cycle, where a proof produced by one curve is +efficiently verified in the circuit of the other curve. However, some of these verifier +checks can actually be efficiently performed in the native circuit; these are "deferred" +to the next native circuit (see diagram below) instead of being immediately passed over to +the other curve. + +![](https://i.imgur.com/l4HrYgE.png) diff --git a/book/src/background/upa.md b/book/src/background/upa.md new file mode 100644 index 0000000..4ffaeb9 --- /dev/null +++ b/book/src/background/upa.md @@ -0,0 +1,77 @@ +# [WIP] UltraPLONK arithmetisation +We work over a multiplicative subgroup +$\mathcal{H} =\{1, \omega, \omega^2, \cdots, \omega^{n-1}\},$ where $\omega$ is primitive +root of unity, in the Lagrange basis corresponding to these points. + +## Polynomial rules +A polynomial rule defines a constraint that must hold between its specified columns at +every row (i.e. at every point in the multiplicative subgroup). + +e.g. + +```text +a * sa + b * sb + a * b * sm + c * sc + PI = 0 +``` + +## Columns +- **fixed (i.e. "selector") columns**: fixed for all instances of a particular circuit. + These columns toggle parts of a polynomial rule "on" or "off" to form a "custom gate". +- **advice columns**: variable values assigned in each instance of the circuit. + Corresponds to the prover's secret witness. +- **public input**: like advice columns, but publicly known values. + +Each column is a vector of $n$ values, e.g. $\mathbf{a} = [a_0, a_1, \cdots, a_{n-1}]$. We +can think of the vector as the evaluation form of the column polynomial +$a(X), X \in \mathcal{H}.$ To recover the coefficient form, we can use +[Lagrange interpolation](polynomials.md#lagrange-interpolation), such that +$a(\omega^i) = a_i.$ + +## Copy constraints +- Define permutation between a set of columns, e.g. $\sigma(a, b, c)$ +- Copy specific cells between these columns, e.g. $b_1 = c_0$ +- Construct permuted columns which should evaluate to same value as original columns + +## Permutation grand product +$$Z(\omega^i) := \prod_{0 \leq j \leq i} \frac{C_k(\omega^j) + \beta\delta^k \omega^j + \gamma}{C_k(\omega^j) + \beta S_k(\omega^j) + \gamma},$$ +where $i = 0, \cdots, n-1$ indexes over the size of the multiplicative subgroup, and +$k = 0, \cdots, m-1$ indexes over the advice columns involved in the permutation. This is +a running product, where each term includes the cumulative product of the terms before it. + +> TODO: what is $\delta$? keep columns linearly independent + +Check the constraints: + +1. First term is equal to one + $$\mathcal{L}_0(X) \cdot (1 - Z(X)) = 0$$ + +2. Running product is well-constructed. For each row, we check that this holds: + $$Z(\omega^i) \cdot{(C(\omega^i) + \beta S_k(\omega^i) + \gamma)} - Z(\omega^{i-1}) \cdot{(C(\omega^i) + \delta^k \beta \omega^i + \gamma)} = 0$$ + Rearranging gives + $$Z(\omega^i) = Z(\omega^{i-1}) \frac{C(\omega^i) + \beta\delta^k \omega^i + \gamma}{C(\omega^i) + \beta S_k(\omega^i) + \gamma},$$ + which is how we defined the grand product polynomial in the first place. + +### Lookup +Reference: [Generic Lookups with PLONK (DRAFT)](/LTPc5f-3S0qNF6MtwD-Tdg?view) + +### Vanishing argument +We want to check that the expressions defined by the gate constraints, permutation +constraints and loookup constraints evaluate to zero at all points in the multiplicative +subgroup. To do this, the prover collapses all the expressions into one polynomial +$$H(X) = \sum_{i=0}^e y^i E_i(X),$$ +where $e$ is the number of expressions and $y$ is a random challenge used to keep the +constraints linearly independent. The prover then divides this by the vanishing polynomial +(see section: [Vanishing polynomial](polynomials.md#vanishing-polynomial)) and commits to +the resulting quotient + +$$\text{Commit}(Q(X)), \text{where } Q(X) = \frac{H(X)}{Z_H(X)}.$$ + +The verifier responds with a random evaluation point $x,$ to which the prover replies with +the claimed evaluations $q = Q(x), \{e_i\}_{i=0}^e = \{E_i(x)\}_{i=0}^e.$ Now, all that +remains for the verifier to check is that the evaluations satisfy + +$$q \stackrel{?}{=} \frac{\sum_{i=0}^e y^i e_i}{Z_H(x)}.$$ + +Notice that we have yet to check that the committed polynomials indeed evaluate to the +claimed values at +$x, q \stackrel{?}{=} Q(x), \{e_i\}_{i=0}^e \stackrel{?}{=} \{E_i(x)\}_{i=0}^e.$ +This check is handled by the polynomial commitment scheme (described in the next section).