close

You are currently browsing the tag archive for the ‘polynomials’ tag.

The notorious Jacobian conjecture can be formulated concretely over the complex numbers as follows.

Conjecture 1 (Jacobian Conjecture) Let {F:{\bf C}^n \rightarrow {\bf C}^n} be a polynomial map in {n} complex variables, whose Jacobian {\mathrm{det} DF} is a non-zero constant. Then {F} is invertible (with polynomial inverse).

The condition that the Jacobian {\mathrm{det} DF} is non-zero is equivalent to {F} being locally invertible. (The implication of local invertibility from non-vanishing Jacobian follows from the inverse function theorem; the converse implication can be derived from the Weierstrass preparation theorem, but is omitted here; see also Lemma 5 of this previous blog post.) Also, from the fundamental theorem of algebra, once the Jacobian polynomial {\mathrm{det} DF} is non-zero, it must be constant. So the hypothesis “Jacobian {\mathrm{det} DF} is a non-zero constant” can be replaced with “{F} is locally invertible”. So the Jacobian conjecture can be viewed as an assertion that local invertibility implies global invertibility. The complex numbers can be easily replaced with other fields of characteristic zero by the Lefschetz principle, but I prefer to work in the concrete setting of the complex numbers.

It was recently shown (using the Fable AI) that the conjecture is false in three dimensions (and thus in higher dimensions as well):

Theorem 2 (Counterexample to conjecture) There exists a polynomial {F : {\bf C}^3 \rightarrow {\bf C}^3} which has non-zero constant Jacobian, but is not invertible.

The conjecture remains open in two dimensions, and is easy to establish in one dimension.

The example can be stated completely explicitly: one can take

\displaystyle  F(z_1,z_2,z_3) = \Big((1+z_1 z_2)^3 z_3 + z_2^2 (1+z_1z_2) (4+3z_1z_2), \ \ \ \ \ (1)

\displaystyle  z_2 + 3 z_1 (1+z_1z_2)^2 z_3 + 3 z_1 z_2^2 (4+3z_1z_2),

\displaystyle 2 z_1 - 3 z_1^2 z_2 - z_1^3 z_3\Big)

and one can verify by a brief calculation that

\displaystyle  \mathrm{det} DF = -2

and

\displaystyle  F(0,0,-1/4) = F(1,-3/2, 13/2) = F(-1,3/2,13/2)

\displaystyle  = (-1/4,0,0).

While this is an extremely quick verification, the construction presented in this fashion appears like a massive miracle. The polynomial {F} has degree seven, so a priori the Jacobian {\mathrm{det} DF} ought to be a polynomial in three variables of degree as large as {3 \times 6 = 18}, so the fact that all non-constant coefficients of this polynomial vanish looks like a massive cancellation involving {\binom{18+3}{3}-1 = 1329} equations, which is much larger than the {3 \times \binom{7+3}{3} = 360} degrees of freedom for a generic degree seven polynomial map of three variables. So finding such a polynomial looks highly unlikely to be located by brute force.

The example has since been retroactively explained in more geometric terms. As a “digestion” exercise to myself, I sought to write this explanation with relatively little use of algebraic geometry, in a manner that minimizes the amount of “miracles” required, although there are still a few places where some remarkable phenomena occur.

It is convenient to use the local injectivity formulation, and to generalize the domain {{\bf C}^3} to an equivalent affine variety. Namely, we will show

Theorem 3 (Counterexample, reformulated) There exists an affine variety {X \subset {\bf C}^5} that is isomorphic to {{\bf C}^3} by polynomial changes of variable, and a polynomial map {F : X \rightarrow {\bf C}^3} which is locally injective, but not globally injective.

Clearly one can get from Theorem 3 to Theorem 2 by composing with the isomorphism {X \cong {\bf C}^3} and using the previously mentioned fact that local injectivity implies non-zero constant Jacobian. Our objective is now to find data {X}, {F : X \rightarrow {\bf C}^3} that obeys three separate properties:

  • (a) {F} is locally injective on {X}.
  • (b) {F} is not globally injective on {X}.
  • (c) {X} is isomorphic to {{\bf C}^3} by polynomial changes of variable.
The advantage of splitting the problem in to these three components is that we can build towards each of them separately.

(A pedantic remark: strictly speaking, in the arguments below, we not only replace the domain {{\bf C}^3} of {F} by an equivalent variety {X}, but also replace the range {{\bf C}^3} of {F} by an equivalent variety {V}. But the equivalence between {V} and {{\bf C}^3} is a boring linear isomorphism ({V} will just be a hyperplane in a four-dimensional vector space {\mathrm{Sym}^3({\bf C}^2)}), so we do not highlight this aspect of the construction.)

It turns out that {F} and {X} can be built out of the operation of multiplication of low degree polynomials. Namely, consider the following three simple affine spaces:

  • The space {\mathrm{Sym}^1({\bf C}^2)} of linear homogeneous polynomials {L(z,w) = az + bw} of two complex variables {z,w}.
  • The space {\mathrm{Sym}^2({\bf C}^2)} of quadratic homogeneous polynomials {Q(z,w) = cz^2 + dzw + ew^2} of two complex variables {z,w}.
  • The space {\mathrm{Sym}^3({\bf C}^2)} of cubic homogeneous polynomials {C(z,w) = fz^3 + gz^2w + hzw^2 + iw^3} of two complex variables {z,w}.
(The notation {\mathrm{Sym}^k(V)} here refers to the {k^{th}} symmetric power of a vector space {V}.) Clearly these spaces are isomorphic to {{\bf C}^2, {\bf C}^3, {\bf C}^4} respectively. Furthermore, we have a multiplication map {F : \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) \rightarrow \mathrm{Sym}^3({\bf C}^2)}, mapping a pair {(L,Q)} of a linear polynomial {L} and a quadratic polynomial {Q} to a cubic polynomial

\displaystyle F(L,Q) := LQ.

(Right now, the domain and range of this map {F} is larger dimensional than the target of three; we will cut the dimensions down to three as the argument progresses.)

The map {F}, essentially a map from {{\bf C}^5} to {{\bf C}^4}, is clearly polynomial; it is given explicitly in coordinates as

\displaystyle  F( (a,b), (c,d,e) ) = (ac, ad + bc, ae + bd, be). \ \ \ \ \ (2)

The map {F} also enjoys two basic (and commuting) symmetries:
  • If one applies a scaling {(L,Q) \mapsto (\lambda_1 L, \lambda_2 Q)} for some non-zero complex numbers {\lambda_1, \lambda_2}, then the product {LQ} is scaled by {C \mapsto \lambda_1 \lambda_2 C}: {F( \lambda_1 L, \lambda_2 Q) = \lambda_1 \lambda_2 F(L,Q)}.
  • If one applies a change of variables {(L, Q) \mapsto (L \circ T, Q \circ T)} for some invertible linear transformation {T \in \mathrm{SL}_2({\bf C})}, then the product {LQ} is transformed by {C \mapsto C \circ T}: {F(L \circ T, Q \circ T) = F(L,Q) \circ T}.
So this map enjoys a huge amount of equivariance, basically with respect to an action of the five-dimensional group {{\bf C}^\times \times {\bf C}^\times \times \mathrm{SL}_2({\bf C})}.

The five-dimensional domain {\mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2)} is of course larger than the four-dimensional range {\mathrm{Sym}^3({\bf C}^2)}, so the map {F} clearly cannot be injective. This can already be seen from the scaling symmetry, as the specific scalings

\displaystyle  (L, Q) \mapsto (\lambda L, \lambda^{-1} Q) \ \ \ \ \ (3)

for {\lambda \in {\bf C}^\times} modify the linear and quadratic polynomials {L,Q} but not their product {C = LQ}. But even if one quotients out by this symmetry (3) to cut the dimension of the domain down to four, the map {F} is still not injective for the following basic reason. A generically chosen cubic polynomial {C} will split into the product {C = L_1 L_2 L_3} of three independent linear polynomials. Then there are three pairs

\displaystyle  (L_1, L_2 L_3), (L_2, L_1 L_3), (L_3, L_1 L_2) \ \ \ \ \ (4)

which all map to the same cubic polynomial

\displaystyle  F(L_1, L_2 L_3) = F(L_2, L_1 L_3) = F(L_3, L_1 L_2) = C

under the multiplication map {F}, but are not related to each other by scaling symmetry (3). Thus, we see that even after quotienting out by the scaling symmetry (3), the multiplication map {F} is generically non-injective in a three-to-one fashion. Thus we already have achieved something resembling goal (b)!

It will be convenient to “spend” the scaling symmetry {(L, Q) \mapsto (\lambda L, \lambda^{-1} Q)} to obtain a useful normalization. If {L(z,w) = az+bw} is a linear polynomial and {Q(z,w) = cz^2 + dzw + ew^2} is a quadratic polynomial, the (homogeneous) resultant {\mathrm{Res}(L,Q)} can be defined by the determinant

\displaystyle  \mathrm{Res}(L,Q) = \begin{vmatrix} a & b & 0 \\ 0 & a & b \\ c & d & e \end{vmatrix} = a^2 e - abd + c b^2. \ \ \ \ \ (5)

If we have a factorization

\displaystyle  L(z,w) = a (z - \alpha w), \quad Q(z,w) = c (z - \beta_1 w)(z - \beta_2 w)

then the resultant can also be described as

\displaystyle  \mathrm{Res}(L,Q) = a^2 c (\alpha - \beta_1) (\alpha - \beta_2).

Thus the resultant measures whether the linear polynomial {L} and the quadratic polynomial {Q} share a common root. A fundamental fact about resultants is that they are {SL_2}-invariant: for any {T \in SL_2({\bf C})}, we have

\displaystyle  \mathrm{Res}(L \circ T, Q \circ T) = \mathrm{Res}(L,Q).

One way to see this is to check it first for translations {(z,w) \mapsto (z + hw, w)} (which translate the roots {\alpha,\beta_1,\beta_2} by {-h} while leaving {a,c} unchanged) and for inversions {(z,w) \mapsto (w,-z)} (which map {\alpha,\beta_1,\beta_2} to {-1/\alpha, -1/\beta_1, -1/\beta_2} while mapping {a,c} to {a\alpha} and {c\beta_1 \beta_2} respectively), and then noting that these transformations generate all of {SL_2({\bf C})}. They also interact very nicely with scaling:

\displaystyle  \mathrm{Res}(\lambda_1 L, \lambda_2 Q) = \lambda_1^2 \lambda_2 \mathrm{Res}(L,Q).

In particular, the scaling symmetry (3) multiplies {\mathrm{Res}(L,Q)} by {\lambda}:

\displaystyle  \mathrm{Res}(\lambda L, \lambda^{-1} Q) = \lambda \mathrm{Res}(L,Q). \ \ \ \ \ (6)

Thus, we can (generically) normalize away this scaling symmetry by imposing the condition

\displaystyle  \mathrm{Res}(L,Q) = 1. \ \ \ \ \ (7)

We now have a restricted multiplication map (which by abuse of notation we will continue to call {F}) from the four-dimensional variety

\displaystyle  \{ (L,Q) \in \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) : \mathrm{Res}(L,Q) = 1\} \ \ \ \ \ (8)

to the four-dimensional space {\mathrm{Sym}^3({\bf C}^2)}. This map {F} is still not globally injective, as we can take the three pairs in (4) from before and apply the scaling (3) separately to each of the three pairs to obtain the normalization (7). So we have kept property (b). Furthermore, this map retains the {SL_2}-equivariance (and also one remaining scaling symmetry, though we will not make much further use of that symmetry).

But we now also have property (a)! Suppose we want to show the local injectivity of {F} in the neighborhood of a pair {(L,Q)} with {\mathrm{Res}(L,Q) = 1}. As the resultant is non-vanishing, the root {\alpha} of {L} (which exists in the Riemann sphere, or projective line if you prefer) is distinct from the two roots {\beta_1, \beta_2} of {Q} (though the latter two roots could be equal to each other). Applying the {SL_2} action (which performs Möbius transforms on the roots), one can assume without loss of generality that {\alpha} is the point at infinity (or equivalently {a=0}), thus {L(z,w) = b w} for some complex number {b} and {Q(z,w) = c (z - \beta_1 w)(z - \beta_2 w)} for some complex numbers {c, \beta_1, \beta_2}, with the resultant condition (7) simplifies to {cb^2 = 1} (so in particular {c,b} are also non-zero). It is then clear that if one perturbs {L} and {Q} by a small amount (say, modifying each coefficient by {O(\varepsilon)}), then the root {\alpha=\infty} of {L} will perturb to something large ({\gg 1/\varepsilon}), while the roots {\beta_1,\beta_2} of {Q} stay bounded. Thus, just from knowledge of the product {F(L,Q)}, one can reconstruct which of the three roots of this cubic polynomial will be the perturbed root of {L}, and which two will be the perturbed roots of {Q}; from this and (6), (7) we can also reconstruct the leading coefficient {c} of {Q}, and this completely determines both {L} and {Q}. This establishes the local injectivity property (a). (In fact it is étale, but we will not need the machinery of étale maps here.)

Unfortunately, (the four-dimensional analogue of) condition (c) fails: the quadric hypersurface (8) is not isomorphic to the affine space {{\bf C}^4}. But we can try to get around this by passing to a three-dimensional slice. Let {V} be some three-dimensional affine plane of {\mathrm{Sym}^3({\bf C}^2)} (which we will take to avoid the origin for technical reasons), then we can restrict {F} as a map from the set

\displaystyle  \{ (L,Q) \in \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) : \mathrm{Res}(L,Q) = 1; \ \ \ \ \ (9)

\displaystyle  F(L,Q) \in V\}

to {V}. {V} is clearly identifiable (by linear changes of coordinate) to {{\bf C}^3}. As {F} was already locally invertible, it remains locally invertible under restriction; and because generic cubic polynomials {C} had three preimages under {F} in (8), this continues to be the case after restricting to (9) (unless {V} was somehow so degenerate that it had no generic elements, but this turns out to be impossible). So we have retained properties (a) and (b). The miracle is that, with a good choice of {V}, we can also obtain (c) and obtain the desired counterexample to the Jacobian conjecture: despite appearances, the variety (9) is in fact equivalent to the affine space {{\bf C}^3} by polynomial changes of variable!

Let’s see how. The affine hyperplanes in {\mathrm{Sym}^3({\bf C}^2)} avoiding the origin are parameterized by the dual space of {\mathrm{Sym}^3({\bf C}^2)} avoiding the origin, which one can think of as the non-zero third order homogeneous differential operators {D = j \partial_z^3 + k \partial_z^2 \partial_w + l \partial_z \partial_w^2 + m \partial_w^3} in two variables. Indeed, every such operator {D} generates an affine hyperplane {\{ C \in \mathrm{Sym}^3({\bf C}^2) : D(C) = 1\}} that avoids the origin, and conversely by duality every affine hyperplane avoiding the origin arises in this form uniquely. Just as the cubic polynomials in {\mathrm{Sym}^3({\bf C}^2)} can be factored into three linear polynomials, the differential operators in the dual space {\mathrm{Sym}^3({\bf C}^2)^*} can also be factored into three linear differential operators, e.g.,

\displaystyle  D = j (\partial_z - \gamma_1 \partial_w) (\partial_z - \gamma_2 \partial_w) (\partial_z - \gamma_3 \partial_w)

in the case that {j} is non-zero. The {SL_2} action moves the roots {\gamma_1,\gamma_2,\gamma_3} around the Riemann sphere by Möbius transformations. As these transformations are {3}-transitive, the actual selection of such roots is not too important (and the scaling symmetry similarly makes the choice of leading coefficient {j} unimportant); the only thing to keep track of is whether the roots repeat. Up to the symmetries, there are in fact just three different equivalence classes of differential operator {D} (and thus of affine hyperplane {V}) to consider:
  • Operators where the three roots {\gamma_1,\gamma_2,\gamma_3} are all distinct, thus {D = D_1 D_2 D_3} for independent first-order operators {D_1,D_2,D_3}.
  • Operators where two roots coincide and one is distinct, thus {D = D_1^2 D_2} for independent first-order operators {D_1,D_2}.
  • Operators where all three roots coincide, thus {D = D_1^3} for some first-order operator {D_1}.

It turns out that the affine miracle for (9) occurs precisely in the second case, when {D} has two identical roots. I do not have a completely satisfactory geometric explanation for this miracle, but one can verify it by the following coordinate computation.

By applying the {SL_2} action, we can normalize so that {D = \frac{1}{2} \partial_z^2 \partial_w}, thus {V} is now the affine hyperplane of cubic polynomials {C(z,w) = f z^3 + g z^2 w + h z w^2 + i w^3} with {g=1}. Using (2) and (5), the variety (9) can now be described explicitly in coordinates as

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1 \}. \ \ \ \ \ (10)

At first glance this seems to be a generic-looking variety cut out by a cubic equation and a quadratic equation – hardly a candidate to be affine! But observe that if {a} is non-zero, then the second equation {ad+bc = 1} can be solved for {d},

\displaystyle  d = \frac{1 - bc}{a} \ \ \ \ \ (11)

and the first equation {a^2 e - abd + cb^2 = 1} can be solved for {e},

\displaystyle  e = \frac{1 + abd - cb^2}{a^2}. \ \ \ \ \ (12)

Putting these two equations together, we see that as long as one removes the case {a=0}, the quintuple {(a,b,c,d,e)} is uniquely determined by {(a,b,c)} by a change of variables which is Laurent in {a} and polynomial in {b,c}. Thus we have a nice birational equivalence

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1; a \neq 0 \}

\displaystyle  \cong \{ (a,b,c) \in {\bf C}^3 : a \neq 0 \}.

Thus we have already almost established property (c): the variety (9) becomes birationally equivalent to {{\bf C}^3} after cutting out the {a=0} subvariety. In particular, for each fixed non-zero value {a_0} of {a}, the corresponding fiber

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1; a = a_0 \}

of (10) is equivalent to {{\bf C}^2} by polynomial changes of variable, since we can reconstruct {d,e} from the coordinates {b,c} by the polynomial formulae

\displaystyle  d = \frac{1-bc}{a_0}; \quad e = \frac{1 + a_0 d b - c b^2}{a_0^2}.

So we just need to glue back in the {a=0} fiber. Indeed, from (10) we see that the fiber at {0} is just

\displaystyle  \{ (0,b,c,d,e) \in {\bf C}^5 : cb^2 = 1; bc = 1 \}.

Now we observe a key miracle: the cubic equation {cb^2 = 1} and quadratic equation {bc = 1} have a unique affine solution {b=c=1} (as opposed to the six possible solutions that Bezout’s theorem might suggest – the other five solutions live on the line at infinity). So the fiber here is also affine:

\displaystyle  \{ (0,1,1,d,e) \in {\bf C}^5 : d, e \in {\bf C} \}.

This is extremely encouraging for the purposes of establishing property (c), as it strongly suggests that the variety (10) has the structure of an {{\bf C}^2}-bundle over {{\bf C}^1}, which is already extremely close to being isomorphic to the affine space {{\bf C}^3}. The main remaining task is to make sure that nothing singular happens in the limit {a \rightarrow 0}, and that a global polynomial coordinate chart for (10) that covers both the {a \neq 0} and {a = 0} fibers can be constructed.

The standard way to proceed here is to manipulate various tangent spaces using the modern machinery of algebraic geometry and commutative algebra, but given my own background, I prefer to adopt the language of analysis, and in particular big-O notation (in place of the ideals used in algebraic geometry), in order to investigate the limit {a \rightarrow 0} by hand. On the variety (10), let us use {O(X)} to denote any multiple of {X} by a polynomial expression in {a,b,c,d,e}. Thus, for instance, the equation {ad + bc = 1} implies that

\displaystyle  bc = 1 + O(a) \ \ \ \ \ (13)

while the equation {a^2 e - abd + cb^2 = 1} implies that

\displaystyle  cb^2 = 1 + O(a) \ \ \ \ \ (14)

as well as the more refined estimate

\displaystyle  cb^2 = 1 + abd + O(a^2). \ \ \ \ \ (15)

In the {a=0} case we could conclude that {b=c=1}. Now we perturb this observation. Multiplying (13) by {b} we have {b^2 c = b + O(a)}, which on substitution into (14) gives {b = 1 + O(a)}; substituting this back into either (13) or (14) also gives {c = 1 + O(a)}.

We can get some more precise asymptotics by also taking advantage of (15). Substituting {ad+bc=1} into (15), we obtain after some algebra

\displaystyle  2 cb^2 = 1 + b + O(a^2).

So if we write {b = 1+O(a)} more explicitly as {b = 1 + a y}, then we have

\displaystyle  2 c (1 + 2ay + O(a^2)) = 2 + ay + O(a^2)

and thus

\displaystyle  c = 1 - \frac{3}{2} ay + O(a^2). \ \ \ \ \ (16)

Substituting this back into (11) gives an asymptotic for {d}:

\displaystyle  d = \frac{1 - bc}{a}

\displaystyle = \frac{1 - (1 + ay) (1 - \frac{3}{2} ay + O(a^2))}{a}

\displaystyle  =\frac{1}{2} y + O(a).

Finally, one can insert these estimates into (12), although one only gets a trivial bound in this case:

\displaystyle  e = \frac{1 + abd - cb^2}{a^2}

\displaystyle = \frac{1 + a (1+O(a)) (\frac{1}{2} y + O(a)) - (1 - \frac{3}{2} ay + O(a^2)) (1 + ay)^2}{a^2}

\displaystyle  = O(1).

Expanding the {O(a^2)} error term in (16) as {a^2 z}, and doing a little more algebra, we thus have a polynomial change of variables

\displaystyle  a = a

\displaystyle  b = 1 + ay

\displaystyle  c = 1 - \frac{3}{2} ay + a^2 z

\displaystyle  d = \frac{1-bc}{a} = \frac{1}{2} y - az + \frac{3}{2} a y^2 - a^2 yz

\displaystyle  e = \frac{1 + abd - cb^2}{a^2} = -2z + 4y^2 - 4ayz + 3ay^3 - 2a^2 y^2 z

which completely parameterizes the variety (10) by polynomial combinations of three coordinates {a,y,z}. This already gives (c) and thus completes the proof of Theorem 3.

The previous computations, when expanded out, also gives polynomial inverse maps:

\displaystyle  a = a

\displaystyle  y = 2bd - ae

\displaystyle  z = 2d^2 + ce + 6bd^2 + 3bce - \frac{9}{2} e

The map from {(a,y,z)} to the {(f,h,i)} coefficients of {F(L,Q)} (dropping the {g} coefficient which is constrained to equal {1}), we obtain a polynomial map

\displaystyle  (a,y,z) \mapsto (G_1(a,y,z), G_2(a,y,z), G_3(a,y,z))

with

\displaystyle  \begin{array}{rl}  G_1(a,y,z) &= a - \frac{3}{2} a^2 y + a^3 z \\ G_2(a,y,z) &= \frac{1}{2} y - 3az + 6ay^2 - 6a^2 yz + \frac{9}{2} a^2 y^3 - 3a^3 y^2 z \\ G_3(a,y,z) &= -2z + 4y^2 - 6ayz + 7ay^3 - 6a^2 y^2 z + 3a^2 y^4 \\ & \quad - 2 a^3 y^3 z \end{array}

which theory predicts to have a constant Jacobian, and indeed one can calculate that the Jacobian is {-1}. This is essentially the original example up to trivial changes of variable; indeed, one can check that the map

\displaystyle  (a, y, -2z) \mapsto (G_3(a,y,z), 2G_2(a,y,z), 2G_1(a,y,z))

is exactly the map {F} given in (1).

AI disclosure: I used an AI chatbot to discuss various aspects of this problem and to confirm several of the calculations made here.

I’ve just uploaded to the arXiv my preprint The maximal length of the Erdős–Herzog–Piranian lemniscate in high degree. This paper resolves (in the asymptotic regime of sufficiently high degree) an old question about the polynomial lemniscates

\displaystyle  \partial E_1(p) := \{ z: |p(z)| = 1\}

attached to monic polynomials {p} of a given degree {n}, and specifically the question of bounding the arclength {\ell(\partial E_1(p))} of such lemniscates. For instance, when {p(z)=z^n}, the lemniscate is the unit circle and the arclength is {2\pi}; this in fact turns out to be the minimum possible length amongst all (connected) lemniscates, a result of Pommerenke. However, the question of the largest lemniscate length is open. The leading candidate for the extremizer is the polynomial

\displaystyle  p_0(z)= z^n-1

whose lemniscate is quite convoluted, with an arclength that can be computed asymptotically as

\displaystyle  \ell(\partial E_1(p_0)) = B(1/2,n/2) = 2n + 4 \log 2 + O(1/n)

where {B} is the beta function.

BERJAYA

(The images here were generated using AlphaEvolve and Gemini.) A reasonably well-known conjecture of Erdős, Herzog, and Piranian (Erdős problem 114) asserts that this is indeed the maximizer, thus {\ell(\partial E_1(p)) \leq \ell(\partial E_1(p_0))} for all monic polynomials of degree {n}.

There have been several partial results towards this conjecture. For instance, Eremenko and Hayman verified the conjecture when {n=2}. Asympotically, bounds of the form {\ell(\partial E_1(p)) \leq (C+o(1)) n} had been known for various {C} such as {\pi}, {2\pi}, or {4\pi}; a significant advance was made by Fryntov and Nazarov, who obtained the asymptotically sharp upper bound

\displaystyle  \ell(\partial E_1(p)) \leq 2n + O(n^{7/8})

and also obtained the sharp conjecture {\ell(\partial E_1(p)) \leq \ell(\partial E_1(p_0))} for {p} sufficiently close to {p_0}. In that paper, the authors comment that the {O(n^{7/8})} error could be improved, but that {O(n^{1/2})} seemed to be the limit of their method.

I recently explored this problem with the optimization tool AlphaEvolve, where I found that when I assigned this tool the task of optimizing {\ell(\partial E_1(p))} for a given degree {n}, that the tool rapidly converged to choosing {p} to be equal to {p_0} (up to the rotation and translation symmetries of the problem). This suggested to me that the conjecture was true for all {n}, though of course this was far from a rigorous proof. AlphaEvolve also provided some useful visualization code for these lemniscates which I have incorporated into the paper (and this blog post), and which helped build my intuition for this problem; I view this sort of “vibe-coded visualization” as another practical use-case of present-day AI tools.

In this paper, we iteratively improve upon the Fryntov-Nazarov method to obtain the following bounds, in increasing order of strength:

  • (i) {\ell(\partial E_1(p)) \leq 2n + O(n^{1/2})}.
  • (ii) {\ell(\partial E_1(p)) \leq 2n + O(1)}.
  • (iii) {\ell(\partial E_1(p)) \leq 2n + 4 \log 2 + o(1)}.
  • (iv) {\ell(\partial E_1(p)) \leq \ell(\partial E_1(p_0))} for sufficiently large {n}.
In particular, the Erdős–Herzog–Piranian conjecture is now verified for sufficiently large {n}.

The proof of these bounds is somewhat circuitious and technical, with the analysis from each part of this result used as a starting point for the next one. For this blog post, I would like to focus on the main ideas of the arguments.

A key difficulty is that there are relatively few tools for upper bounding the arclength of a curve; indeed, the coastline paradox already shows that curves can have infinite length even when bounded. Thus, one needs to utilize some smooth or algebraic structure on the curve to hope for good upper bounds. One possible approach is via the Crofton formula, using Bezout’s theorem to control the intersection of the curve with various lines. This is already good enough to get bounds of the form {\ell(\partial E_1(p)) \leq O(n)} (for instance by combining it with other known tools to control the diameter of the lemniscate), but it seems challenging to use this approach to get bounds close to the optimal {\sim 2n}.

Instead, we follow Fryntov–Nazarov and utilize Stokes’ theorem to convert the arclength into an area integral. A typical identity used in that paper is

\displaystyle  \ell(\partial E_1(p)) = 2 \int_{E_1(p)} |p'|\ dA - \int_{E_1(p)} \frac{|p\varphi|}{\varphi} \psi\ dA

where {dA} is area measure, {\varphi = p'/p} is the log-derivative of {p}, {\psi = p''/p'} is the log-derivative of {p'}, and {E_1(p)} is the region {\{ |p| \leq 1 \}}. This and the triangle inequality already lets one prove bounds of the form {\ell(\partial E_1(p)) \leq 2\pi n + O(n^{1/2})} by controlling {\int_{E_1(p)} |\psi|\ dA} using the triangle inequality, the Hardy-Littlewood rearrangement inequality, and the Polya capacity inequality.

But this argument does not fully capture the oscillating nature of the phase {\frac{|\varphi|}{\varphi}} on one hand, and the oscillating nature of {E_1(p)} on the other. Fryntov–Nazarov exploited these oscillations with some additional decompositions and integration by parts arguments. By optimizing these arguments, I was able to establish an inequality of the form

\displaystyle  \ell(\partial E_1(p)) \leq \frac{1}{\pi} \int_{E_2(p)} |\psi|\ dA + O(X_1+X_2+X_3+X_4+X_5) \ \ \ \ \ (1)

where {E_2(p) = \{ |p| \leq 2\}} is an enlargement of {E_1(p)} (that is significantly less oscillatory, as displayed in the figure below), and {X_1,\dots,X_5} are certain error terms that can be controlled by a number of standard tools (e.g., the Grönwall area theorem).

One can heuristically justify (1) as follows. Suppose we work in a region where the functions {\psi}, {p} are roughly constant: {\psi(z) \approx \psi_0}, {p(z) \approx c_0}. For simplicity let us normalize {\psi_0} to be real, and {c_0} to be negative real. In order to have a non-trivial lemniscate in this region, {c_0} should be close to {-1}. Because the unit circle {\partial D(0,1)} is tangent to the line {\{ w: \Re w = -1\}} at {-1}, the lemniscate condition {|p(z)|=1} is then heuristically approximated by the condition that {\Re (p(z) + 1) = 0}. On the other hand, the hypothesis {\psi(z) \approx \psi_0} suggests that {p'(z) \approx A e^{\psi_0 z}} for some amplitude {z}, which heuristically integrates to {p(z) \approx c_0 + \frac{A}{\psi_0} e^{\psi_0 z}}. Writing {\frac{A}{\psi_0}} in polar coordinates as {Re^{i\theta}} and {z} in Cartesian coordinates as {x+iy}, the condition {\Re(p(z)+1)=0} can then be rearranged after some algebra as

\displaystyle  \cos (\psi_0 y + \theta) \approx -\frac{1+c_0}{Re^{\psi_0 x}}.

If the right-hand side is much larger than {1} in magnitude, we thus expect the lemniscate to be empty in this region; but if instead the right-hand side is much less than {1} in magnitude, we expect the lemniscate to behave like a periodic sequence of horizontal lines of spacing {\frac{\pi}{\psi_0}}. This makes the main terms on both sides of (1) approximately agree (per unit area).

A graphic illustration of (1) (provided by Gemini) is shown below, where the dark spots correspond to small values of {\psi} that act to “repel” (and shorten) the lemniscate. (The bright spots correspond to the critical points of {p}, which in this case consist of six critical points at the origin and one at both of {+1/2} and {-1/2}.)

BERJAYA

By choosing parameters appropriately, one can show that {X_1,\dots,X_5 = O(\sqrt{n})} and {\frac{1}{\pi} \int_{E_2(p)} |\psi|\ dA \leq 2n + O(1)}, yielding the first bound {\ell(\partial E_1(p)) \leq 2n + O(n^{1/2})}. However, by a more careful inspection of the arguments, and in particular measuring the defect in the triangle inequality

\displaystyle  |\psi(z)| = |\sum_\zeta \frac{1}{z-\zeta}| \leq \sum_\zeta \frac{1}{|z-\zeta|},

where {\zeta} ranges over critical points. From some elementary geometry, one can show that the more the critical points {\zeta} are dispersed away from each other (in an {\ell^1} sense), the more one can gain over the triangle inequality here; conversely, the {L^1} dispersion {\sum_\zeta |\zeta|} of the critical points (after normalizing so that these critical points have mean zero) can be used to improve the control on the error terms {X_1,\dots,X_5}. Optimizing this strategy leads to the second bound

\displaystyle \ell(\partial E_1(p)) \leq 2n + O(1).

At this point, the only remaining cases that need to be handled are the ones with bounded dispersion: {\sum_\zeta |\zeta| \leq 1}. In this case, one can do some elementary manipulations of the factorization

\displaystyle  p'(z) = n \prod_\zeta (z-\zeta)

to obtain some quite precise control on the asymptotics of {p(z)} and {p'(z)}; for instance, we will be able to obtain an approximation of the form

\displaystyle  p(z) \approx -1 + \frac{z p'(z)}{n}

with high accuracy, as long as {z} is not too close to the origin or to critical points. This, combined with direct arclength computations, can eventually lead to the third estimate

\displaystyle \ell(\partial E_1(p)) \leq 2n + 4 \log 2 + o(1).

The last remaining cases to handle are those of small dispersion, {\sum_\zeta |\zeta| = o(1)}. An extremely careful version of the previous analysis can now give an estimate of the shape

\displaystyle \ell(\partial E_1(p)) \leq \ell(\partial E_1(p_0)) - c \|p\| + o(\|p\|)

for an absolute constant {c>0}, where {\|p\|} is a measure of how close {p} is to {p_0} (it is equal to the dispersion {\sum_\zeta |\zeta|} plus an additional term {n |1 + p(0)|^{1/n}} to deal with the constant term {p(0)}). This establishes the final bound (for {n} large enough), and even shows that the only extremizer is {p_0} (up to translation and rotation symmetry).

As someone who had a relatively light graduate education in algebra, the import of Yoneda’s lemma in category theory has always eluded me somewhat; the statement and proof are simple enough, but definitely have the “abstract nonsense” flavor that one often ascribes to this part of mathematics, and I struggled to connect it to the more grounded forms of intuition, such as those based on concrete examples, that I was more comfortable with. There is a popular MathOverflow post devoted to this question, with many answers that were helpful to me, but I still felt vaguely dissatisfied. However, recently when pondering the very concrete concept of a polynomial, I managed to accidentally stumble upon a special case of Yoneda’s lemma in action, which clarified this lemma conceptually for me. In the end it was a very simple observation (and would be extremely pedestrian to anyone who works in an algebraic field of mathematics), but as I found this helpful to a non-algebraist such as myself, and I thought I would share it here in case others similarly find it helpful.

In algebra we see a distinction between a polynomial form (also known as a formal polynomial), and a polynomial function, although this distinction is often elided in more concrete applications. A polynomial form in, say, one variable with integer coefficients, is a formal expression {P} of the form

\displaystyle  P = a_d {\mathrm n}^d + \dots + a_1 {\mathrm n} + a_0 \ \ \ \ \ (1)

where {a_0,\dots,a_d} are coefficients in the integers, and {{\mathrm n}} is an indeterminate: a symbol that is often intended to be interpreted as an integer, real number, complex number, or element of some more general ring {R}, but is for now a purely formal object. The collection of such polynomial forms is denoted {{\bf Z}[{\mathrm n}]}, and is a commutative ring.

A polynomial form {P} can be interpreted in any ring {R} (even non-commutative ones) to create a polynomial function {P_R : R \rightarrow R}, defined by the formula

\displaystyle  P_R(n) := a_d n^d + \dots + a_1 n + a_0 \ \ \ \ \ (2)

for any {n \in R}. This definition (2) looks so similar to the definition (1) that we usually abuse notation and conflate {P} with {P_R}. This conflation is supported by the identity theorem for polynomials, that asserts that if two polynomial forms {P, Q} agree at an infinite number of (say) complex numbers, thus {P_{\bf C}(z) = Q_{\bf C}(z)} for infinitely many {z}, then they agree {P=Q} as polynomial forms (i.e., their coefficients match). But this conflation is sometimes dangerous, particularly when working in finite characteristic. For instance:

  • (i) The linear forms {{\mathrm n}} and {-{\mathrm n}} are distinct as polynomial forms, but agree when interpreted in the ring {{\bf Z}/2{\bf Z}}, since {n = -n} for all {n \in {\bf Z}/2{\bf Z}}.
  • (ii) Similarly, if {p} is a prime, then the degree one form {{\mathrm n}} and the degree {p} form {{\mathrm n}^p} are distinct as polynomial forms (and in particular have distinct degrees), but agree when interpreted in the ring {{\bf Z}/p{\bf Z}}, thanks to Fermat’s little theorem.
  • (iii) The polynomial form {{\mathrm n}^2+1} has no roots when interpreted in the reals {{\bf R}}, but has roots when interpreted in the complex numbers {{\bf C}}. Similarly, the linear form {2{\mathrm n}-1} has no roots when interpreted in the integers {{\bf Z}}, but has roots when interpreted in the rationals {{\bf Q}}.

The above examples show that if one only interprets polynomial forms in a specific ring {R}, then some information about the polynomial could be lost (and some features of the polynomial, such as roots, may be “invisible” to that interpretation). But this turns out not to be the case if one considers interpretations in all rings simultaneously, as we shall now discuss.

If {R, S} are two different rings, then the polynomial functions {P_R: R \rightarrow R} and {P_S: S \rightarrow S} arising from interpreting a polynomial form {P} in these two rings are, strictly speaking, different functions. However, they are often closely related to each other. For instance, if {R} is a subring of {S}, then {P_R} agrees with the restriction of {P_S} to {R}. More generally, if there is a ring homomorphism {\phi: R \rightarrow S} from {R} to {S}, then {P_R} and {P_S} are intertwined by the relation

\displaystyle  \phi \circ P_R = P_S \circ \phi, \ \ \ \ \ (3)

which basically asserts that ring homomorphism respect polynomial operations. Note that the previous observation corresponded to the case when {\phi} was an inclusion homomorphism. Another example comes from the complex conjugation automorphism {z \mapsto \overline{z}} on the complex numbers, in which case (3) asserts the identity

\displaystyle  \overline{P_{\bf C}(z)} = P_{\bf C}(\overline{z})

for any polynomial function {P_{\bf C}} on the complex numbers, and any complex number {z}.

What was surprising to me (as someone who had not internalized the Yoneda lemma) was that the converse statement was true: if one had a function {F_R: R \rightarrow R} associated to every ring {R} that obeyed the intertwining relation

\displaystyle  \phi \circ F_R = F_S \circ \phi \ \ \ \ \ (4)

for every ring homomorphism {\phi: R \rightarrow S}, then there was a unique polynomial form {P \in {\bf Z}[\mathrm{n}]} such that {F_R = P_R} for all rings {R}. This seemed surprising to me because the functions {F} were a priori arbitrary functions, and as an analyst I would not expect them to have polynomial structure. But the fact that (4) holds for all rings {R,S} and all homomorphisms {\phi} is in fact rather powerful. As an analyst, I am tempted to proceed by first working with the ring {{\bf C}} of complex numbers and taking advantage of the aforementioned identity theorem, but this turns out to be tricky because {{\bf C}} does not “talk” to all the other rings {R} enough, in the sense that there are not always as many ring homomorphisms from {{\bf C}} to {R} as one would like. But there is in fact a more elementary argument that takes advantage of a particularly relevant (and “talkative”) ring to the theory of polynomials, namely the ring {{\bf Z}[\mathrm{n}]} of polynomials themselves. Given any other ring {R}, and any element {n} of that ring, there is a unique ring homomorphism {\phi_{R,n}: {\bf Z}[\mathrm{n}] \rightarrow R} from {{\bf Z}[\mathrm{n}]} to {R} that maps {\mathrm{n}} to {n}, namely the evaluation map

\displaystyle  \phi_{R,n} \colon a_d {\mathrm n}^d + \dots + a_1 {\mathrm n} + a_0 \mapsto a_d n^d + \dots + a_1 n + a_0

that sends a polynomial form to its evaluation at {n}. Applying (4) to this ring homomorphism, and specializing to the element {\mathrm{n}} of {{\bf Z}[\mathrm{n}]}, we conclude that

\displaystyle  \phi_{R,n}( F_{{\bf Z}[\mathrm{n}]}(\mathrm{n}) ) = F_R( n )

for any ring {R} and any {n \in R}. If we then define {P \in {\bf Z}[\mathrm{n}]} to be the formal polynomial

\displaystyle  P := F_{{\bf Z}[\mathrm{n}]}(\mathrm{n}),

then this identity can be rewritten as

\displaystyle  F_R = P_R

and so we have indeed shown that the family {F_R} arises from a polynomial form {P}. Conversely, from the identity

\displaystyle  P = P_{{\bf Z}[\mathrm{n}]}(\mathrm{n})

valid for any polynomial form {P}, we see that two polynomial forms {P,Q} can only generate the same polynomial functions {P_R, Q_R} for all rings {R} if they are identical as polynomial forms. So the polynomial form {P} associated to the family {F_R} is unique.

We have thus created an identification of form and function: polynomial forms {P} are in one-to-one correspondence with families of functions {F_R} obeying the intertwining relation (4). But this identification can be interpreted as a special case of the Yoneda lemma, as follows. There are two categories in play here: the category {\mathbf{Ring}} of rings (where the morphisms are ring homomorphisms), and the category {\mathrm{Set}} of sets (where the morphisms are arbitrary functions). There is an obvious forgetful functor {\mathrm{Forget}: \mathbf{Ring} \rightarrow \mathbf{Set}} between these two categories that takes a ring and removes all of the algebraic structure, leaving behind just the underlying set. A collection {F_R: R \rightarrow R} of functions (i.e., {\mathbf{Set}}-morphisms) for each {R} in {\mathbf{Ring}} that obeys the intertwining relation (4) is precisely the same thing as a natural transformation from the forgetful functor {\mathrm{Forget}} to itself. So we have identified formal polynomials in {{\bf Z}[\mathbf{n}]} as a set with natural endomorphisms of the forgetful functor:

\displaystyle  \mathrm{Forget}({\bf Z}[\mathbf{n}]) \equiv \mathrm{Hom}( \mathrm{Forget}, \mathrm{Forget} ). \ \ \ \ \ (5)

Informally: polynomial forms are precisely those operations on rings that are respected by ring homomorphisms.

What does this have to do with Yoneda’s lemma? Well, remember that every element {n} of a ring {R} came with an evaluation homomorphism {\phi_{R,n}: {\bf Z}[\mathrm{n}] \rightarrow R}. Conversely, every homomorphism from {{\bf Z}[\mathrm{n}]} to {R} will be of the form {\phi_{R,n}} for a unique {n} – indeed, {n} will just be the image of {\mathrm{n}} under this homomorphism. So the evaluation homomorphism provides a one-to-one correspondence between elements of {R}, and ring homomorphisms in {\mathrm{Hom}({\bf Z}[\mathrm{n}], R)}. This correspondence is at the level of sets, so this gives the identification

\displaystyle  \mathrm{Forget} \equiv \mathrm{Hom}({\bf Z}[\mathrm{n}], -).

Thus our identification can be written as

\displaystyle  \mathrm{Forget}({\bf Z}[\mathbf{n}]) \equiv \mathrm{Hom}( \mathrm{Hom}({\bf Z}[\mathrm{n}], -), \mathrm{Forget} )

which is now clearly a special case of the Yoneda lemma

\displaystyle  F(A) \equiv \mathrm{Hom}( \mathrm{Hom}(A, -), F )

that applies to any functor {F: {\mathcal C} \rightarrow \mathbf{Set}} from a (locally small) category {{\mathcal C}} and any object {A} in {{\mathcal C}}. And indeed if one inspects the standard proof of this lemma, it is essentially the same argument as the argument we used above to establish the identification (5). More generally, it seems to me that the Yoneda lemma is often used to identify “formal” objects with their “functional” interpretations, as long as one simultaneously considers interpretations across an entire category (such as the category of rings), as opposed to just a single interpretation in a single object of the category in which there may be some loss of information due to the peculiarities of that specific object. Grothendieck’s “functor of points” interpretation of a scheme, discussed in this previous blog post, is one typical example of this.

I’ve just uploaded to the arXiv my paper “Sendov’s conjecture for sufficiently high degree polynomials“. This paper is a contribution to an old conjecture of Sendov on the zeroes of polynomials:

Conjecture 1 (Sendov’s conjecture) Let {f: {\bf C} \rightarrow {\bf C}} be a polynomial of degree {n \geq 2} that has all zeroes in the closed unit disk {\{ z: |z| \leq 1 \}}. If {\lambda_0} is one of these zeroes, then {f'} has at least one zero in {\{z: |z-\lambda_0| \leq 1\}}.

It is common in the literature on this problem to normalise {f} to be monic, and to rotate the zero {\lambda_0} to be an element {a} of the unit interval {[0,1]}. As it turns out, the location of {a} on this unit interval {[0,1]} ends up playing an important role in the arguments.

Many cases of this conjecture are already known, for instance

In particular, in high degrees the only cases left uncovered by prior results are when {a} is close (but not too close) to {0}, or when {a} is close (but not too close) to {1}; see Figure 1 of my paper.

Our main result covers the high degree case uniformly for all values of {a \in [0,1]}:

Theorem 2 There exists an absolute constant {n_0} such that Sendov’s conjecture holds for all {n \geq n_0}.

In principle, this reduces the verification of Sendov’s conjecture to a finite time computation, although our arguments use compactness methods and thus do not easily provide an explicit value of {n_0}. I believe that the compactness arguments can be replaced with quantitative substitutes that provide an explicit {n_0}, but the value of {n_0} produced is likely to be extremely large (certainly much larger than {9}).

Because of the previous results (particularly those of Chalebgwa and Chijiwa), we will only need to establish the following two subcases of the above theorem:

Theorem 3 (Sendov’s conjecture near the origin) Under the additional hypothesis {a = o(1/\log n)}, Sendov’s conjecture holds for sufficiently large {n}.

Theorem 4 (Sendov’s conjecture near the unit circle) Under the additional hypothesis {1-o(1) \leq a \leq 1 - \varepsilon_0^n} for a fixed {\varepsilon_0>0}, Sendov’s conjecture holds for sufficiently large {n}.

We approach these theorems using the “compactness and contradiction” strategy, assuming that there is a sequence of counterexamples whose degrees {n} going to infinity, using various compactness theorems to extract various asymptotic objects in the limit {n \rightarrow \infty}, and somehow using these objects to derive a contradiction. There are many ways to effect such a strategy; we will use a formalism that I call “cheap nonstandard analysis” and which is common in the PDE literature, in which one repeatedly passes to subsequences as necessary whenever one invokes a compactness theorem to create a limit object. However, the particular choice of asymptotic formalism one selects is not of essential importance for the arguments.

I also found it useful to use the language of probability theory. Given a putative counterexample {f} to Sendov’s conjecture, let {\lambda} be a zero of {f} (chosen uniformly at random among the {n} zeroes of {f}, counting multiplicity), and let {\zeta} similarly be a uniformly random zero of {f'}. We introduce the logarithmic potentials

\displaystyle  U_\lambda(z) := {\bf E} \log \frac{1}{|z-\lambda|}; \quad U_\zeta(z) := {\bf E} \log \frac{1}{|z-\zeta|}

and the Stieltjes transforms

\displaystyle  s_\lambda(z) := {\bf E} \frac{1}{z-\lambda}; \quad s_\zeta(z) := {\bf E} \log \frac{1}{z-\zeta}.

Standard calculations using the fundamental theorem of algebra yield the basic identities

\displaystyle  U_\lambda(z) = \frac{1}{n} \log \frac{1}{|f(z)|}; \quad U_\zeta(z) = \frac{1}{n-1} \log \frac{n}{|f'(z)|}

and

\displaystyle  s_\lambda(z) = \frac{1}{n} \frac{f'(z)}{f(z)}; \quad s_\zeta(z) = \frac{1}{n-1} \frac{f''(z)}{f'(z)} \ \ \ \ \ (1)

and in particular the random variables {\lambda, \zeta} are linked to each other by the identity

\displaystyle  U_\lambda(z) - \frac{n-1}{n} U_\zeta(z) = \frac{1}{n} \log |s_\lambda(z)|. \ \ \ \ \ (2)

On the other hand, the hypotheses of Sendov’s conjecture (and the Gauss-Lucas theorem) place {\lambda,\zeta} inside the unit disk {\{ z:|z| \leq 1\}}. Applying Prokhorov’s theorem, and passing to a subsequence, one can then assume that the random variables {\lambda,\zeta} converge in distribution to some limiting random variables {\lambda^{(\infty)}, \zeta^{(\infty)}} (possibly defined on a different probability space than the original variables {\lambda,\zeta}), also living almost surely inside the unit disk. Standard potential theory then gives the convergence

\displaystyle  U_\lambda(z) \rightarrow U_{\lambda^{(\infty)}}(z); \quad U_\zeta(z) \rightarrow U_{\zeta^{(\infty)}}(z) \ \ \ \ \ (3)

and

\displaystyle  s_\lambda(z) \rightarrow s_{\lambda^{(\infty)}}(z); \quad s_\zeta(z) \rightarrow s_{\zeta^{(\infty)}}(z) \ \ \ \ \ (4)

at least in the local {L^1} sense. Among other things, we then conclude from the identity (2) and some elementary inequalities that

\displaystyle  U_{\lambda^{(\infty)}}(z) = U_{\zeta^{(\infty)}}(z)

for all {|z|>1}. This turns out to have an appealing interpretation in terms of Brownian motion: if one takes two Brownian motions in the complex plane, one originating from {\lambda^{(\infty)}} and one originating from {\zeta^{(\infty)}}, then the location where these Brownian motions first exit the unit disk {\{ z: |z| \leq 1 \}} will have the same distribution. (In our paper we actually replace Brownian motion with the closely related formalism of balayage.) This turns out to connect the random variables {\lambda^{(\infty)}}, {\zeta^{(\infty)}} quite closely to each other. In particular, with this observation and some additional arguments involving both the unique continuation property for harmonic functions and Grace’s theorem (discussed in this previous post), with the latter drawn from the prior work of Dégot, we can get very good control on these distributions:

Theorem 5
  • (i) If {a = o(1)}, then {\lambda^{(\infty)}, \zeta^{(\infty)}} almost surely lie in the semicircle {\{ e^{i\theta}: \pi/2 \leq \theta \leq 3\pi/2\}} and have the same distribution.
  • (ii) If {a = 1-o(1)}, then {\lambda^{(\infty)}} is uniformly distributed on the circle {\{ z: |z|=1\}}, and {\zeta^{(\infty)}} is almost surely zero.

In case (i) (and strengthening the hypothesis {a=o(1)} to {a=o(1/\log n)} to control some technical contributions of “outlier” zeroes of {f}), we can use this information about {\lambda^{(\infty)}} and (4) to ensure that the normalised logarithmic derivative {\frac{1}{n} \frac{f'}{f} = s_\lambda} has a non-negative winding number in a certain small (but not too small) circle around the origin, which by the argument principle is inconsistent with the hypothesis that {f} has a zero at {a = o(1)} and that {f'} has no zeroes near {a}. This is how we establish Theorem 3.

Case (ii) turns out to be more delicate. This is because there are a number of “near-counterexamples” to Sendov’s conjecture that are compatible with the hypotheses and conclusion of case (ii). The simplest such example is {f(z) = z^n - 1}, where the zeroes {\lambda} of {f} are uniformly distributed amongst the {n^{th}} roots of unity (including at {a=1}), and the zeroes of {f'} are all located at the origin. In my paper I also discuss a variant of this construction, in which {f'} has zeroes mostly near the origin, but also acquires a bounded number of zeroes at various locations {\lambda_1+o(1),\dots,\lambda_m+o(1)} inside the unit disk. Specifically, we take

\displaystyle  f(z) := \left(z + \frac{c_2}{n}\right)^{n-m} P(z) - \left(a + \frac{c_2}{n}\right)^{n-m} P(a)

where {a = 1 - \frac{c_1}{n}} for some constants {0 < c_1 < c_2} and

\displaystyle  P(z) := (z-\lambda_1) \dots (z-\lambda_m).

By a perturbative analysis to locate the zeroes of {f}, one eventually would be able to arrive at a true counterexample to Sendov’s conjecture if these locations {\lambda_1,\dots,\lambda_m} were in the open lune

\displaystyle  \{ \lambda: |\lambda| < 1 < |\lambda-1| \}

and if one had the inequality

\displaystyle  c_2 - c_1 - c_2 \cos \theta + \sum_{j=1}^m \log \left|\frac{1 - \lambda_j}{e^{i\theta} - \lambda_j}\right| < 0 \ \ \ \ \ (5)

for all {0 \leq \theta \leq 2\pi}. However, if one takes the mean of this inequality in {\theta}, one arrives at the inequality

\displaystyle  c_2 - c_1 + \sum_{j=1}^m \log |1 - \lambda_j| < 0

which is incompatible with the hypotheses {c_2 > c_1} and {|\lambda_j-1| > 1}. In order to extend this argument to more general polynomials {f}, we require a stability analysis of the endpoint equation

\displaystyle  c_2 - c_1 + c_2 \cos \theta + \sum_{j=1}^m \log \left|\frac{1 - \lambda_j}{e^{i\theta} - \lambda_j}\right| = 0 \ \ \ \ \ (6)

where we now only assume the closed conditions {c_2 \geq c_1} and {|\lambda_j-1| \geq 1}. The above discussion then places all the zeros {\lambda_j} on the arc

\displaystyle  \{ \lambda: |\lambda| < 1 = |\lambda-1|\} \ \ \ \ \ (7)

and if one also takes the second Fourier coefficient of (6) one also obtains the vanishing second moment

\displaystyle  \sum_{j=1}^m \lambda_j^2 = 0.

These two conditions are incompatible with each other (except in the degenerate case when all the {\lambda_j} vanish), because all the non-zero elements {\lambda} of the arc (7) have argument in {\pm [\pi/3,\pi/2]}, so in particular their square {\lambda^2} will have negative real part. It turns out that one can adapt this argument to the more general potential counterexamples to Sendov’s conjecture (in the form of Theorem 4). The starting point is to use (1), (4), and Theorem 5(ii) to obtain good control on {f''/f'}, which one then integrates and exponentiates to get good control on {f'}, and then on a second integration one gets enough information about {f} to pin down the location of its zeroes to high accuracy. The constraint that these zeroes lie inside the unit disk then gives an inequality resembling (5), and an adaptation of the above stability analysis is then enough to conclude. The arguments here are inspired by the previous arguments of Miller, which treated the case when {a} was extremely close to {1} via a similar perturbative analysis; the main novelty is to control the error terms not in terms of the magnitude of the largest zero {\zeta} of {f'} (which is difficult to manage when {n} gets large), but rather by the variance of those zeroes, which ends up being a more tractable expression to keep track of.

[UPDATE, Feb 1, 2021: the strategy sketched out below has been successfully implemented to rigorously obtain the desired implication in this recent preprint of Giulio Bresciani.]
I recently came across this question on MathOverflow asking if there are any polynomials {P} of two variables with rational coefficients, such that the map {P: {\bf Q} \times {\bf Q} \rightarrow {\bf Q}} is a bijection. The answer to this question is almost surely “no”, but it is remarkable how hard this problem resists any attempt at rigorous proof. (MathOverflow users with enough privileges to see deleted answers will find that there are no less than seventeen deleted attempts at a proof in response to this question!)
On the other hand, the one surviving response to the question does point out this paper of Poonen which shows that assuming a powerful conjecture in Diophantine geometry known as the Bombieri-Lang conjecture (discussed in this previous post), it is at least possible to exhibit polynomials {P: {\bf Q} \times {\bf Q} \rightarrow {\bf Q}} which are injective.
I believe that it should be possible to also rule out the existence of bijective polynomials {P: {\bf Q} \times {\bf Q} \rightarrow {\bf Q}} if one assumes the Bombieri-Lang conjecture, and have sketched out a strategy to do so, but filling in the gaps requires a fair bit more algebraic geometry than I am capable of. So as a sort of experiment, I would like to see if a rigorous implication of this form (similarly to the rigorous implication of the Erdos-Ulam conjecture from the Bombieri-Lang conjecture in my previous post) can be crowdsourced, in the spirit of the polymath projects (though I feel that this particular problem should be significantly quicker to resolve than a typical such project).
Here is how I imagine a Bombieri-Lang-powered resolution of this question should proceed (modulo a large number of unjustified and somewhat vague steps that I believe to be true but have not established rigorously). Suppose for contradiction that we have a bijective polynomial {P: {\bf Q} \times {\bf Q} \rightarrow {\bf Q}}. Then for any polynomial {Q: {\bf Q} \rightarrow {\bf Q}} of one variable, the surface

\displaystyle  S_Q := \{ (x,y,z) \in \mathbb{A}^3: P(x,y) = Q(z) \}

has infinitely many rational points; indeed, every rational {z \in {\bf Q}} lifts to exactly one rational point in {S_Q}. I believe that for “typical” {Q} this surface {S_Q} should be irreducible. One can now split into two cases:

  • (a) The rational points in {S_Q} are Zariski dense in {S_Q}.
  • (b) The rational points in {S_Q} are not Zariski dense in {S_Q}.

Consider case (b) first. By definition, this case asserts that the rational points in {S_Q} are contained in a finite number of algebraic curves. By Faltings’ theorem (a special case of the Bombieri-Lang conjecture), any curve of genus two or higher only contains a finite number of rational points. So all but finitely many of the rational points in {S_Q} are contained in a finite union of genus zero and genus one curves. I think all genus zero curves are birational to a line, and all the genus one curves are birational to an elliptic curve (though I don’t have an immediate reference for this). These curves {C} all can have an infinity of rational points, but very few of them should have “enough” rational points {C \cap {\bf Q}^3} that their projection {\pi(C \cap {\bf Q}^3) := \{ z \in {\bf Q} : (x,y,z) \in C \hbox{ for some } x,y \in {\bf Q} \}} to the third coordinate is “large”. In particular, I believe

  • (i) If {C \subset {\mathbb A}^3} is birational to an elliptic curve, then the number of elements of {\pi(C \cap {\bf Q}^3)} of height at most {H} should grow at most polylogarithmically in {H} (i.e., be of order {O( \log^{O(1)} H )}.
  • (ii) If {C \subset {\mathbb A}^3} is birational to a line but not of the form {\{ (f(z), g(z), z) \}} for some rational {f,g}, then then the number of elements of {\pi(C \cap {\bf Q}^3)} of height at most {H} should grow slower than {H^2} (in fact I think it can only grow like {O(H)}).

I do not have proofs of these results (though I think something similar to (i) can be found in Knapp’s book, and (ii) should basically follow by using a rational parameterisation {\{(f(t),g(t),h(t))\}} of {C} with {h} nonlinear). Assuming these assertions, this would mean that there is a curve of the form {\{ (f(z),g(z),z)\}} that captures a “positive fraction” of the rational points of {S_Q}, as measured by restricting the height of the third coordinate {z} to lie below a large threshold {H}, computing density, and sending {H} to infinity (taking a limit superior). I believe this forces an identity of the form

\displaystyle  P(f(z), g(z)) = Q(z) \ \ \ \ \ (1)

for all {z}. Such identities are certainly possible for some choices of {Q} (e.g. {Q(z) = P(F(z), G(z))} for arbitrary polynomials {F,G} of one variable) but I believe that the only way that such identities hold for a “positive fraction” of {Q} (as measured using height as before) is if there is in fact a rational identity of the form

\displaystyle  P( f_0(z), g_0(z) ) = z

for some rational functions {f_0,g_0} with rational coefficients (in which case we would have {f = f_0 \circ Q} and {g = g_0 \circ Q}). But such an identity would contradict the hypothesis that {P} is bijective, since one can take a rational point {(x,y)} outside of the curve {\{ (f_0(z), g_0(z)): z \in {\bf Q} \}}, and set {z := P(x,y)}, in which case we have {P(x,y) = P(f_0(z), g_0(z) )} violating the injective nature of {P}. Thus, modulo a lot of steps that have not been fully justified, we have ruled out the scenario in which case (b) holds for a “positive fraction” of {Q}.
This leaves the scenario in which case (a) holds for a “positive fraction” of {Q}. Assuming the Bombieri-Lang conjecture, this implies that for such {Q}, any resolution of singularities of {S_Q} fails to be of general type. I would imagine that this places some very strong constraints on {P,Q}, since I would expect the equation {P(x,y) = Q(z)} to describe a surface of general type for “generic” choices of {P,Q} (after resolving singularities). However, I do not have a good set of techniques for detecting whether a given surface is of general type or not. Presumably one should proceed by viewing the surface {\{ (x,y,z): P(x,y) = Q(z) \}} as a fibre product of the simpler surface {\{ (x,y,w): P(x,y) = w \}} and the curve {\{ (z,w): Q(z) = w \}} over the line {\{w \}}. In any event, I believe the way to handle (a) is to show that the failure of general type of {S_Q} implies some strong algebraic constraint between {P} and {Q} (something in the spirit of (1), perhaps), and then use this constraint to rule out the bijectivity of {P} by some further ad hoc method.

(This post is mostly intended for my own reference, as I found myself repeatedly looking up several conversions between polynomial bases on various occasions.)

Let {\mathrm{Poly}_{\leq n}} denote the vector space of polynomials {P:{\bf R} \rightarrow {\bf R}} of one variable {x} with real coefficients of degree at most {n}. This is a vector space of dimension {n+1}, and the sequence of these spaces form a filtration:

\displaystyle  \mathrm{Poly}_{\leq 0} \subset \mathrm{Poly}_{\leq 1} \subset \mathrm{Poly}_{\leq 2} \subset \dots

A standard basis for these vector spaces are given by the monomials {x^0, x^1, x^2, \dots}: every polynomial {P(x)} in {\mathrm{Poly}_{\leq n}} can be expressed uniquely as a linear combination of the first {n+1} monomials {x^0, x^1, \dots, x^n}. More generally, if one has any sequence {Q_0(x), Q_1(x), Q_2(x)} of polynomials, with each {Q_n} of degree exactly {n}, then an easy induction shows that {Q_0(x),\dots,Q_n(x)} forms a basis for {\mathrm{Poly}_{\leq n}}.

In particular, if we have two such sequences {Q_0(x), Q_1(x), Q_2(x),\dots} and {R_0(x), R_1(x), R_2(x), \dots} of polynomials, with each {Q_n} of degree {n} and each {R_k} of degree {k}, then {Q_n} must be expressible uniquely as a linear combination of the polynomials {R_0,R_1,\dots,R_n}, thus we have an identity of the form

\displaystyle  Q_n(x) = \sum_{k=0}^n c_{QR}(n,k) R_k(x)

for some change of basis coefficients {c_{QR}(n,k) \in {\bf R}}. These coefficients describe how to convert a polynomial expressed in the {Q_n} basis into a polynomial expressed in the {R_k} basis.

Many standard combinatorial quantities {c(n,k)} involving two natural numbers {0 \leq k \leq n} can be interpreted as such change of basis coefficients. The most familiar example are the binomial coefficients {\binom{n}{k}}, which measures the conversion from the shifted monomial basis {(x+1)^n} to the monomial basis {x^k}, thanks to (a special case of) the binomial formula:

\displaystyle  (x+1)^n = \sum_{k=0}^n \binom{n}{k} x^k,

thus for instance

\displaystyle  (x+1)^3 = \binom{3}{0} x^0 + \binom{3}{1} x^1 + \binom{3}{2} x^2 + \binom{3}{3} x^3

\displaystyle  = 1 + 3x + 3x^2 + x^3.

More generally, for any shift {h}, the conversion from {(x+h)^n} to {x^k} is measured by the coefficients {h^{n-k} \binom{n}{k}}, thanks to the general case of the binomial formula.

But there are other bases of interest too. For instance if one uses the falling factorial basis

\displaystyle  (x)_n := x (x-1) \dots (x-n+1)

then the conversion from falling factorials to monomials is given by the Stirling numbers of the first kind {s(n,k)}:

\displaystyle  (x)_n = \sum_{k=0}^n s(n,k) x^k,

thus for instance

\displaystyle  (x)_3 = s(3,0) x^0 + s(3,1) x^1 + s(3,2) x^2 + s(3,3) x^3

\displaystyle  = 0 + 2 x - 3x^2 + x^3

and the conversion back is given by the Stirling numbers of the second kind {S(n,k)}:

\displaystyle  x^n = \sum_{k=0}^n S(n,k) (x)_k

thus for instance

\displaystyle  x^3 = S(3,0) (x)_0 + S(3,1) (x)_1 + S(3,2) (x)_2 + S(3,3) (x)_3

\displaystyle  = 0 + x + 3 x(x-1) + x(x-1)(x-2).

If one uses the binomial functions {\binom{x}{n} = \frac{1}{n!} (x)_n} as a basis instead of the falling factorials, one of course can rewrite these conversions as

\displaystyle  \binom{x}{n} = \sum_{k=0}^n \frac{1}{n!} s(n,k) x^k

and

\displaystyle  x^n = \sum_{k=0}^n k! S(n,k) \binom{x}{k}

thus for instance

\displaystyle  \binom{x}{3} = 0 + \frac{1}{3} x - \frac{1}{2} x^2 + \frac{1}{6} x^3

and

\displaystyle  x^3 = 0 + \binom{x}{1} + 6 \binom{x}{2} + 6 \binom{x}{3}.

As a slight variant, if one instead uses rising factorials

\displaystyle  (x)^n := x (x+1) \dots (x+n-1)

then the conversion to monomials yields the unsigned Stirling numbers {|s(n,k)|} of the first kind:

\displaystyle  (x)^n = \sum_{k=0}^n |s(n,k)| x^k

thus for instance

\displaystyle  (x)^3 = 0 + 2x + 3x^2 + x^3.

One final basis comes from the polylogarithm functions

\displaystyle  \mathrm{Li}_{-n}(x) := \sum_{j=1}^\infty j^n x^j.

For instance one has

\displaystyle  \mathrm{Li}_1(x) = -\log(1-x)

\displaystyle  \mathrm{Li}_0(x) = \frac{x}{1-x}

\displaystyle  \mathrm{Li}_{-1}(x) = \frac{x}{(1-x)^2}

\displaystyle  \mathrm{Li}_{-2}(x) = \frac{x}{(1-x)^3} (1+x)

\displaystyle  \mathrm{Li}_{-3}(x) = \frac{x}{(1-x)^4} (1+4x+x^2)

\displaystyle  \mathrm{Li}_{-4}(x) = \frac{x}{(1-x)^5} (1+11x+11x^2+x^3)

and more generally one has

\displaystyle  \mathrm{Li}_{-n-1}(x) = \frac{x}{(1-x)^{n+2}} E_n(x)

for all natural numbers {n} and some polynomial {E_n} of degree {n} (the Eulerian polynomials), which when converted to the monomial basis yields the (shifted) Eulerian numbers

\displaystyle  E_n(x) = \sum_{k=0}^n A(n+1,k) x^k.

For instance

\displaystyle  E_3(x) = A(4,0) x^0 + A(4,1) x^1 + A(4,2) x^2 + A(4,3) x^3

\displaystyle  = 1 + 11x + 11x^2 + x^3.

These particular coefficients also have useful combinatorial interpretations. For instance:

  • The binomial coefficient {\binom{n}{k}} is of course the number of {k}-element subsets of {\{1,\dots,n\}}.
  • The unsigned Stirling numbers {|s(n,k)|} of the first kind are the number of permutations of {\{1,\dots,n\}} with exactly {k} cycles. The signed Stirling numbers {s(n,k)} are then given by the formula {s(n,k) = (-1)^{n-k} |s(n,k)|}.
  • The Stirling numbers {S(n,k)} of the second kind are the number of ways to partition {\{1,\dots,n\}} into {k} non-empty subsets.
  • The Eulerian numbers {A(n,k)} are the number of permutations of {\{1,\dots,n\}} with exactly {k} ascents.

These coefficients behave similarly to each other in several ways. For instance, the binomial coefficients {\binom{n}{k}} obey the well known Pascal identity

\displaystyle  \binom{n+1}{k} = \binom{n}{k} + \binom{n}{k-1}

(with the convention that {\binom{n}{k}} vanishes outside of the range {0 \leq k \leq n}). In a similar spirit, the unsigned Stirling numbers {|s(n,k)|} of the first kind obey the identity

\displaystyle  |s(n+1,k)| = n |s(n,k)| + |s(n,k-1)|

and the signed counterparts {s(n,k)} obey the identity

\displaystyle  s(n+1,k) = -n s(n,k) + s(n,k-1).

The Stirling numbers of the second kind {S(n,k)} obey the identity

\displaystyle  S(n+1,k) = k S(n,k) + S(n,k-1)

and the Eulerian numbers {A(n,k)} obey the identity

\displaystyle  A(n+1,k) = (k+1) A(n,k) + (n-k+1) A(n,k-1).

This is a sequel to this previous blog post, in which we discussed the effect of the heat flow evolution

\displaystyle  \partial_t P(t,z) = \partial_{zz} P(t,z)

on the zeroes of a time-dependent family of polynomials {z \mapsto P(t,z)}, with a particular focus on the case when the polynomials {z \mapsto P(t,z)} had real zeroes. Here (inspired by some discussions I had during a recent conference on the Riemann hypothesis in Bristol) we record the analogous theory in which the polynomials instead have zeroes on a circle {\{ z: |z| = \sqrt{q} \}}, with the heat flow slightly adjusted to compensate for this. As we shall discuss shortly, a key example of this situation arises when {P} is the numerator of the zeta function of a curve.

More precisely, let {g} be a natural number. We will say that a polynomial

\displaystyle  P(z) = \sum_{j=0}^{2g} a_j z^j

of degree {2g} (so that {a_{2g} \neq 0}) obeys the functional equation if the {a_j} are all real and

\displaystyle  a_j = q^{g-j} a_{2g-j}

for all {j=0,\dots,2g}, thus

\displaystyle  P(\overline{z}) = \overline{P(z)}

and

\displaystyle  P(q/z) = q^g z^{-2g} P(z)

for all non-zero {z}. This means that the {2g} zeroes {\alpha_1,\dots,\alpha_{2g}} of {P(z)} (counting multiplicity) lie in {{\bf C} \backslash \{0\}} and are symmetric with respect to complex conjugation {z \mapsto \overline{z}} and inversion {z \mapsto q/z} across the circle {\{ |z| = \sqrt{q}\}}. We say that this polynomial obeys the Riemann hypothesis if all of its zeroes actually lie on the circle {\{ z = \sqrt{q}\}}. For instance, in the {g=1} case, the polynomial {z^2 - a_1 z + q} obeys the Riemann hypothesis if and only if {|a_1| \leq 2\sqrt{q}}.

Such polynomials arise in number theory as follows: if {C} is a projective curve of genus {g} over a finite field {\mathbf{F}_q}, then, as famously proven by Weil, the associated local zeta function {\zeta_{C,q}(z)} (as defined for instance in this previous blog post) is known to take the form

\displaystyle  \zeta_{C,q}(z) = \frac{P(z)}{(1-z)(1-qz)}

where {P} is a degree {2g} polynomial obeying both the functional equation and the Riemann hypothesis. In the case that {C} is an elliptic curve, then {g=1} and {P} takes the form {P(z) = z^2 - a_1 z + q}, where {a_1} is the number of {{\bf F}_q}-points of {C} minus {q+1}. The Riemann hypothesis in this case is a famous result of Hasse.

Another key example of such polynomials arise from rescaled characteristic polynomials

\displaystyle  P(z) := \det( 1 - \sqrt{q} F ) \ \ \ \ \ (1)

of {2g \times 2g} matrices {F} in the compact symplectic group {Sp(g)}. These polynomials obey both the functional equation and the Riemann hypothesis. The Sato-Tate conjecture (in higher genus) asserts, roughly speaking, that “typical” polyomials {P} arising from the number theoretic situation above are distributed like the rescaled characteristic polynomials (1), where {F} is drawn uniformly from {Sp(g)} with Haar measure.

Given a polynomial {z \mapsto P(0,z)} of degree {2g} with coefficients

\displaystyle  P(0,z) = \sum_{j=0}^{2g} a_j(0) z^j,

we can evolve it in time by the formula

\displaystyle  P(t,z) = \sum_{j=0}^{2g} \exp( t(j-g)^2 ) a_j(0) z^j,

thus {a_j(t) = \exp(t(j-g)) a_j(0)} for {t \in {\bf R}}. Informally, as one increases {t}, this evolution accentuates the effect of the extreme monomials, particularly, {z^0} and {z^{2g}} at the expense of the intermediate monomials such as {z^g}, and conversely as one decreases {t}. This family of polynomials obeys the heat-type equation

\displaystyle  \partial_t P(t,z) = (z \partial_z - g)^2 P(t,z). \ \ \ \ \ (2)

In view of the results of Marcus, Spielman, and Srivastava, it is also very likely that one can interpret this flow in terms of expected characteristic polynomials involving conjugation over the compact symplectic group {Sp(n)}, and should also be tied to some sort of “{\beta=\infty}” version of Brownian motion on this group, but we have not attempted to work this connection out in detail.

It is clear that if {z \mapsto P(0,z)} obeys the functional equation, then so does {z \mapsto P(t,z)} for any other time {t}. Now we investigate the evolution of the zeroes. Suppose at some time {t_0} that the zeroes {\alpha_1(t_0),\dots,\alpha_{2g}(t_0)} of {z \mapsto P(t_0,z)} are distinct, then

\displaystyle  P(t_0,z) = a_{2g}(0) \exp( t_0g^2 ) \prod_{j=1}^{2g} (z - \alpha_j(t_0) ).

From the inverse function theorem we see that for times {t} sufficiently close to {t_0}, the zeroes {\alpha_1(t),\dots,\alpha_{2g}(t)} of {z \mapsto P(t,z)} continue to be distinct (and vary smoothly in {t}), with

\displaystyle  P(t,z) = a_{2g}(0) \exp( t g^2 ) \prod_{j=1}^{2g} (z - \alpha_j(t) ).

Differentiating this at any {z} not equal to any of the {\alpha_j(t)}, we obtain

\displaystyle  \partial_t P(t,z) = P(t,z) ( g^2 - \sum_{j=1}^{2g} \frac{\alpha'_j(t)}{z - \alpha_j(t)})

and

\displaystyle  \partial_z P(t,z) = P(t,z) ( \sum_{j=1}^{2g} \frac{1}{z - \alpha_j(t)})

and

\displaystyle  \partial_{zz} P(t,z) = P(t,z) ( \sum_{1 \leq j,k \leq 2g: j \neq k} \frac{1}{(z - \alpha_j(t))(z - \alpha_k(t))}).

Inserting these formulae into (2) (expanding {(z \partial_z - g)^2} as {z^2 \partial_{zz} - (2g-1) z \partial_z + g^2}) and canceling some terms, we conclude that

\displaystyle  - \sum_{j=1}^{2g} \frac{\alpha'_j(t)}{z - \alpha_j(t)} = z^2 \sum_{1 \leq j,k \leq 2g: j \neq k} \frac{1}{(z - \alpha_j(t))(z - \alpha_k(t))}

\displaystyle  - (2g-1) z \sum_{j=1}^{2g} \frac{1}{z - \alpha_j(t)}

for {t} sufficiently close to {t_0}, and {z} not equal to {\alpha_1(t),\dots,\alpha_{2g}(t)}. Extracting the residue at {z = \alpha_j(t)}, we conclude that

\displaystyle  - \alpha'_j(t) = 2 \alpha_j(t)^2 \sum_{1 \leq k \leq 2g: k \neq j} \frac{1}{\alpha_j(t) - \alpha_k(t)} - (2g-1) \alpha_j(t)

which we can rearrange as

\displaystyle  \frac{\alpha'_j(t)}{\alpha_j(t)} = - \sum_{1 \leq k \leq 2g: k \neq j} \frac{\alpha_j(t)+\alpha_k(t)}{\alpha_j(t)-\alpha_k(t)}.

If we make the change of variables {\alpha_j(t) = \sqrt{q} e^{i\theta_j(t)}} (noting that one can make {\theta_j} depend smoothly on {t} for {t} sufficiently close to {t_0}), this becomes

\displaystyle  \partial_t \theta_j(t) = \sum_{1 \leq k \leq 2g: k \neq j} \cot \frac{\theta_j(t) - \theta_k(t)}{2}. \ \ \ \ \ (3)

Intuitively, this equation asserts that the phases {\theta_j} repel each other if they are real (and attract each other if their difference is imaginary). If {z \mapsto P(t_0,z)} obeys the Riemann hypothesis, then the {\theta_j} are all real at time {t_0}, then the Picard uniqueness theorem (applied to {\theta_j(t)} and its complex conjugate) then shows that the {\theta_j} are also real for {t} sufficiently close to {t_0}. If we then define the entropy functional

\displaystyle  H(\theta_1,\dots,\theta_{2g}) := \sum_{1 \leq j < k \leq 2g} \log \frac{1}{|\sin \frac{\theta_j-\theta_k}{2}| }

then the above equation becomes a gradient flow

\displaystyle  \partial_t \theta_j(t) = - 2 \frac{\partial H}{\partial \theta_j}( \theta_1(t),\dots,\theta_{2g}(t) )

which implies in particular that {H(\theta_1(t),\dots,\theta_{2g}(t))} is non-increasing in time. This shows that as one evolves time forward from {t_0}, there is a uniform lower bound on the separation between the phases {\theta_1(t),\dots,\theta_{2g}(t)}, and hence the equation can be solved indefinitely; in particular, {z \mapsto P(t,z)} obeys the Riemann hypothesis for all {t > t_0} if it does so at time {t_0}. Our argument here assumed that the zeroes of {z \mapsto P(t_0,z)} were simple, but this assumption can be removed by the usual limiting argument.

For any polynomial {z \mapsto P(0,z)} obeying the functional equation, the rescaled polynomials {z \mapsto e^{-g^2 t} P(t,z)} converge locally uniformly to {a_{2g}(0) (z^{2g} + q^g)} as {t \rightarrow +\infty}. By Rouche’s theorem, we conclude that the zeroes of {z \mapsto P(t,z)} converge to the equally spaced points {\{ e^{2\pi i(j+1/2)/2g}: j=1,\dots,2g\}} on the circle {\{ |z| = \sqrt{q}\}}. Together with the symmetry properties of the zeroes, this implies in particular that {z \mapsto P(t,z)} obeys the Riemann hypothesis for all sufficiently large positive {t}. In the opposite direction, when {t \rightarrow -\infty}, the polynomials {z \mapsto P(t,z)} converge locally uniformly to {a_g(0) z^g}, so if {a_g(0) \neq 0}, {g} of the zeroes converge to the origin and the other {g} converge to infinity. In particular, {z \mapsto P(t,z)} fails the Riemann hypothesis for sufficiently large negative {t}. Thus (if {a_g(0) \neq 0}), there must exist a real number {\Lambda}, which we call the de Bruijn-Newman constant of the original polynomial {z \mapsto P(0,z)}, such that {z \mapsto P(t,z)} obeys the Riemann hypothesis for {t \geq \Lambda} and fails the Riemann hypothesis for {t < \Lambda}. The situation is a bit more complicated if {a_g(0)} vanishes; if {k} is the first natural number such that {a_{g+k}(0)} (or equivalently, {a_{g-j}(0)}) does not vanish, then by the above arguments one finds in the limit {t \rightarrow -\infty} that {g-k} of the zeroes go to the origin, {g-k} go to infinity, and the remaining {2k} zeroes converge to the equally spaced points {\{ e^{2\pi i(j+1/2)/2k}: j=1,\dots,2k\}}. In this case the de Bruijn-Newman constant remains finite except in the degenerate case {k=g}, in which case {\Lambda = -\infty}.

For instance, consider the case when {g=1} and {P(0,z) = z^2 - a_1 z + q} for some real {a_1} with {|a_1| \leq 2\sqrt{q}}. Then the quadratic polynomial

\displaystyle  P(t,z) = e^t z^2 - a_1 z + e^t q

has zeroes

\displaystyle  \frac{a_1 \pm \sqrt{a_1^2 - 4 e^{2t} q}}{2e^t}

and one easily checks that these zeroes lie on the circle {\{ |z|=\sqrt{q}\}} when {t \geq \log \frac{|a_1|}{2\sqrt{q}}}, and are on the real axis otherwise. Thus in this case we have {\Lambda = \log \frac{|a_1|}{2\sqrt{q}}} (with {\Lambda=-\infty} if {a_1=0}). Note how as {t} increases to {+\infty}, the zeroes repel each other and eventually converge to {\pm i \sqrt{q}}, while as {t} decreases to {-\infty}, the zeroes collide and then separate on the real axis, with one zero going to the origin and the other to infinity.

The arguments in my paper with Brad Rodgers (discussed in this previous post) indicate that for a “typical” polynomial {P} of degree {g} that obeys the Riemann hypothesis, the expected time to relaxation to equilibrium (in which the zeroes are equally spaced) should be comparable to {1/g}, basically because the average spacing is {1/g} and hence by (3) the typical velocity of the zeroes should be comparable to {g}, and the diameter of the unit circle is comparable to {1}, thus requiring time comparable to {1/g} to reach equilibrium. Taking contrapositives, this suggests that the de Bruijn-Newman constant {\Lambda} should typically take on values comparable to {-1/g} (since typically one would not expect the initial configuration of zeroes to be close to evenly spaced). I have not attempted to formalise or prove this claim, but presumably one could do some numerics (perhaps using some of the examples of {P} given previously) to explore this further.

Let {P(z) = z^n + a_{n-1} z^{n-1} + \dots + a_0} be a monic polynomial of degree {n} with complex coefficients. Then by the fundamental theorem of algebra, we can factor {P} as

\displaystyle  P(z) = (z-z_1) \dots (z-z_n) \ \ \ \ \ (1)

for some complex zeroes {z_1,\dots,z_n} (possibly with repetition).

Now suppose we evolve {P} with respect to time by heat flow, creating a function {P(t,z)} of two variables with given initial data {P(0,z) = P(z)} for which

\displaystyle  \partial_t P(t,z) = \partial_{zz} P(t,z). \ \ \ \ \ (2)

On the space of polynomials of degree at most {n}, the operator {\partial_{zz}} is nilpotent, and one can solve this equation explicitly both forwards and backwards in time by the Taylor series

\displaystyle  P(t,z) = \sum_{j=0}^\infty \frac{t^j}{j!} \partial_z^{2j} P(0,z).

For instance, if one starts with a quadratic {P(0,z) = z^2 + bz + c}, then the polynomial evolves by the formula

\displaystyle  P(t,z) = z^2 + bz + (c+2t).

As the polynomial {P(t)} evolves in time, the zeroes {z_1(t),\dots,z_n(t)} evolve also. Assuming for sake of discussion that the zeroes are simple, the inverse function theorem tells us that the zeroes will (locally, at least) evolve smoothly in time. What are the dynamics of this evolution?

For instance, in the quadratic case, the quadratic formula tells us that the zeroes are

\displaystyle  z_1(t) = \frac{-b + \sqrt{b^2 - 4(c+2t)}}{2}

and

\displaystyle  z_2(t) = \frac{-b - \sqrt{b^2 - 4(c+2t)}}{2}

after arbitrarily choosing a branch of the square root. If {b,c} are real and the discriminant {b^2 - 4c} is initially positive, we see that we start with two real zeroes centred around {-b/2}, which then approach each other until time {t = \frac{b^2-4c}{8}}, at which point the roots collide and then move off from each other in an imaginary direction.

In the general case, we can obtain the equations of motion by implicitly differentiating the defining equation

\displaystyle  P( t, z_i(t) ) = 0

in time using (2) to obtain

\displaystyle  \partial_{zz} P( t, z_i(t) ) + \partial_t z_i(t) \partial_z P(t,z_i(t)) = 0.

To simplify notation we drop the explicit dependence on time, thus

\displaystyle  \partial_{zz} P(z_i) + (\partial_t z_i) \partial_z P(z_i)= 0.

From (1) and the product rule, we see that

\displaystyle  \partial_z P( z_i ) = \prod_{j:j \neq i} (z_i - z_j)

and

\displaystyle  \partial_{zz} P( z_i ) = 2 \sum_{k:k \neq i} \prod_{j:j \neq i,k} (z_i - z_j)

(where all indices are understood to range over {1,\dots,n}) leading to the equations of motion

\displaystyle  \partial_t z_i = \sum_{k:k \neq i} \frac{2}{z_k - z_i}, \ \ \ \ \ (3)

at least when one avoids those times in which there is a repeated zero. In the case when the zeroes {z_i} are real, each term {\frac{2}{z_k-z_i}} represents a (first-order) attraction in the dynamics between {z_i} and {z_k}, but the dynamics are more complicated for complex zeroes (e.g. purely imaginary zeroes will experience repulsion rather than attraction, as one already sees in the quadratic example). Curiously, this system resembles that of Dyson brownian motion (except with the brownian motion part removed, and time reversed). I learned of the connection between the ODE (3) and the heat equation from this paper of Csordas, Smith, and Varga, but perhaps it has been mentioned in earlier literature as well.

One interesting consequence of these equations is that if the zeroes are real at some time, then they will stay real as long as the zeroes do not collide. Let us now restrict attention to the case of real simple zeroes, in which case we will rename the zeroes as {x_i} instead of {z_i}, and order them as {x_1 < \dots < x_n}. The evolution

\displaystyle  \partial_t x_i = \sum_{k:k \neq i} \frac{2}{x_k - x_i}

can now be thought of as reverse gradient flow for the “entropy”

\displaystyle  H := -\sum_{i,j: i \neq j} \log |x_i - x_j|,

(which is also essentially the logarithm of the discriminant of the polynomial) since we have

\displaystyle  \partial_t x_i = \frac{\partial H}{\partial x_i}.

In particular, we have the monotonicity formula

\displaystyle  \partial_t H = 4E

where {E} is the “energy”

\displaystyle  E := \frac{1}{4} \sum_i (\frac{\partial H}{\partial x_i})^2

\displaystyle  = \sum_i (\sum_{k:k \neq i} \frac{1}{x_k-x_i})^2

\displaystyle  = \sum_{i,k: i \neq k} \frac{1}{(x_k-x_i)^2} + 2 \sum_{i,j,k: i,j,k \hbox{ distinct}} \frac{1}{(x_k-x_i)(x_j-x_i)}

\displaystyle  = \sum_{i,k: i \neq k} \frac{1}{(x_k-x_i)^2}

where in the last line we use the antisymmetrisation identity

\displaystyle  \frac{1}{(x_k-x_i)(x_j-x_i)} + \frac{1}{(x_i-x_j)(x_k-x_j)} + \frac{1}{(x_j-x_k)(x_i-x_k)} = 0.

Among other things, this shows that as one goes backwards in time, the entropy decreases, and so no collisions can occur to the past, only in the future, which is of course consistent with the attractive nature of the dynamics. As {H} is a convex function of the positions {x_1,\dots,x_n}, one expects {H} to also evolve in a convex manner in time, that is to say the energy {E} should be increasing. This is indeed the case:

Exercise 1 Show that

\displaystyle  \partial_t E = 2 \sum_{i,j: i \neq j} (\frac{2}{(x_i-x_j)^2} - \sum_{k: i,j,k \hbox{ distinct}} \frac{1}{(x_k-x_i)(x_k-x_j)})^2.

Symmetric polynomials of the zeroes are polynomial functions of the coefficients and should thus evolve in a polynomial fashion. One can compute this explicitly in simple cases. For instance, the center of mass is an invariant:

\displaystyle  \partial_t \frac{1}{n} \sum_i x_i = 0.

The variance decreases linearly:

Exercise 2 Establish the virial identity

\displaystyle  \partial_t \sum_{i,j} (x_i-x_j)^2 = - 4n^2(n-1).

As the variance (which is proportional to {\sum_{i,j} (x_i-x_j)^2}) cannot become negative, this identity shows that “finite time blowup” must occur – that the zeroes must collide at or before the time {\frac{1}{4n^2(n-1)} \sum_{i,j} (x_i-x_j)^2}.

Exercise 3 Show that the Stieltjes transform

\displaystyle  s(t,z) = \sum_i \frac{1}{x_i - z}

solves the viscous Burgers equation

\displaystyle  \partial_t s = \partial_{zz} s - 2 s \partial_z s,

either by using the original heat equation (2) and the identity {s = - \partial_z P / P}, or else by using the equations of motion (3). This relation between the Burgers equation and the heat equation is known as the Cole-Hopf transformation.

The paper of Csordas, Smith, and Varga mentioned previously gives some other bounds on the lifespan of the dynamics; roughly speaking, they show that if there is one pair of zeroes that are much closer to each other than to the other zeroes then they must collide in a short amount of time (unless there is a collision occuring even earlier at some other location). Their argument extends also to situations where there are an infinite number of zeroes, which they apply to get new results on Newman’s conjecture in analytic number theory. I would be curious to know of further places in the literature where this dynamics has been studied.

The equidistribution theorem asserts that if {\alpha \in {\bf R}/{\bf Z}} is an irrational phase, then the sequence {(n\alpha)_{n=1}^\infty} is equidistributed on the unit circle, or equivalently that

\displaystyle \frac{1}{N} \sum_{n=1}^N F(n\alpha) \rightarrow \int_{{\bf R}/{\bf Z}} F(x)\ dx

for any continuous (or equivalently, for any smooth) function {F: {\bf R}/{\bf Z} \rightarrow {\bf C}}. By approximating {F} uniformly by a Fourier series, this claim is equivalent to that of showing that

\displaystyle \frac{1}{N} \sum_{n=1}^N e(hn\alpha) \rightarrow 0

for any non-zero integer {h} (where {e(x) := e^{2\pi i x}}), which is easily verified from the irrationality of {\alpha} and the geometric series formula. Conversely, if {\alpha} is rational, then clearly {\frac{1}{N} \sum_{n=1}^N e(hn\alpha)} fails to go to zero when {h} is a multiple of the denominator of {\alpha}.

One can then ask for more quantitative information about the decay of exponential sums of {\frac{1}{N} \sum_{n=1}^N e(n \alpha)}, or more generally on exponential sums of the form {\frac{1}{|Q|} \sum_{n \in Q} e(P(n))} for an arithmetic progression {Q} (in this post all progressions are understood to be finite) and a polynomial {P: Q \rightarrow \/{\bf Z}}. It will be convenient to phrase such information in the form of an inverse theorem, describing those phases for which the exponential sum is large. Indeed, we have

Lemma 1 (Geometric series formula, inverse form) Let {Q \subset {\bf Z}} be an arithmetic progression of length at most {N} for some {N \geq 1}, and let {P(n) = n \alpha + \beta} be a linear polynomial for some {\alpha,\beta \in {\bf R}/{\bf Z}}. If

\displaystyle \frac{1}{N} |\sum_{n \in Q} e(P(n))| \geq \delta

for some {\delta > 0}, then there exists a subprogression {Q'} of {Q} of size {|Q'| \gg \delta^2 N} such that {P(n)} varies by at most {\delta} on {Q'} (that is to say, {P(n)} lies in a subinterval of {{\bf R}/{\bf Z}} of length at most {\delta}).

Proof: By a linear change of variable we may assume that {Q} is of the form {\{0,\dots,N'-1\}} for some {N' \geq 1}. We may of course assume that {\alpha} is non-zero in {{\bf R}/{\bf Z}}, so that {\|\alpha\|_{{\bf R}/{\bf Z}} > 0} ({\|x\|_{{\bf R}/{\bf Z}}} denotes the distance from {x} to the nearest integer). From the geometric series formula we see that

\displaystyle |\sum_{n \in Q} e(P(n))| \leq \frac{2}{|e(\alpha) - 1|} \ll \frac{1}{\|\alpha\|_{{\bf R}/{\bf Z}}},

and so {\|\alpha\|_{{\bf R}/{\bf Z}} \ll \frac{1}{\delta N}}. Setting {Q' := \{ n \in Q: n \leq c \delta^2 N \}} for some sufficiently small absolute constant {c}, we obtain the claim. \Box

Thus, in order for a linear phase {P(n)} to fail to be equidistributed on some long progression {Q}, {P} must in fact be almost constant on large piece of {Q}.

As is well known, this phenomenon generalises to higher order polynomials. To achieve this, we need two elementary additional lemmas. The first relates the exponential sums of {P} to the exponential sums of its “first derivatives” {n \mapsto P(n+h)-P(n)}.

Lemma 2 (Van der Corput lemma, inverse form) Let {Q \subset {\bf Z}} be an arithmetic progression of length at most {N}, and let {P: Q \rightarrow {\bf R}/{\bf Z}} be an arbitrary function such that

\displaystyle \frac{1}{N} |\sum_{n \in Q} e(P(n))| \geq \delta \ \ \ \ \ (1)

 

for some {\delta > 0}. Then, for {\gg \delta^2 N} integers {h \in Q-Q}, there exists a subprogression {Q_h} of {Q}, of the same spacing as {Q}, such that

\displaystyle \frac{1}{N} |\sum_{n \in Q_h} e(P(n+h)-P(n))| \gg \delta^2. \ \ \ \ \ (2)

 

Proof: Squaring (1), we see that

\displaystyle \sum_{n,n' \in Q} e(P(n') - P(n)) \geq \delta^2 N^2.

We write {n' = n+h} and conclude that

\displaystyle \sum_{h \in Q-Q} \sum_{n \in Q_h} e( P(n+h)-P(n) ) \geq \delta^2 N^2

where {Q_h := Q \cap (Q-h)} is a subprogression of {Q} of the same spacing. Since {\sum_{n \in Q_h} e( P(n+h)-P(n) ) = O(N)}, we conclude that

\displaystyle |\sum_{n \in Q_h} e( P(n+h)-P(n) )| \gg \delta^2 N

for {\gg \delta^2 N} values of {h} (this can be seen, much like the pigeonhole principle, by arguing via contradiction for a suitable choice of implied constants). The claim follows. \Box

The second lemma (which we recycle from this previous blog post) is a variant of the equidistribution theorem.

Lemma 3 (Vinogradov lemma) Let {I \subset [-N,N] \cap {\bf Z}} be an interval for some {N \geq 1}, and let {\theta \in{\bf R}/{\bf Z}} be such that {\|n\theta\|_{{\bf R}/{\bf Z}} \leq \varepsilon} for at least {\delta N} values of {n \in I}, for some {0 < \varepsilon, \delta < 1}. Then either

\displaystyle N < \frac{2}{\delta}

or

\displaystyle \varepsilon > 10^{-2} \delta

or else there is a natural number {q \leq 2/\delta} such that

\displaystyle \| q \theta \|_{{\bf R}/{\bf Z}} \ll \frac{\varepsilon}{\delta N}.

Proof: We may assume that {N \geq \frac{2}{\delta}} and {\varepsilon \leq 10^{-2} \delta}, since we are done otherwise. Then there are at least two {n \in I} with {\|n \theta \|_{{\bf R}/{\bf Z}} \leq \varepsilon}, and by the pigeonhole principle we can find {n_1 < n_2} in {Q} with {\|n_1 \theta \|_{{\bf R}/{\bf Z}}, \|n_2 \theta \|_{{\bf R}/{\bf Z}} \leq \varepsilon} and {n_2-n_1 \leq \frac{2}{\delta}}. By the triangle inequality, we conclude that there exists at least one natural number {q \leq \frac{2}{\delta}} for which

\displaystyle \| q \theta \|_{{\bf R}/{\bf Z}} \leq 2\varepsilon.

We take {q} to be minimal amongst all such natural numbers, then we see that there exists {a} coprime to {q} and {|\kappa| \leq 2\varepsilon} such that

\displaystyle \theta = \frac{a}{q} + \frac{\kappa}{q}. \ \ \ \ \ (3)

 

If {\kappa=0} then we are done, so suppose that {\kappa \neq 0}. Suppose that {n < m} are elements of {I} such that {\|n\theta \|_{{\bf R}/{\bf Z}}, \|m\theta \|_{{\bf R}/{\bf Z}} \leq \varepsilon} and {m-n \leq \frac{1}{10 \kappa}}. Writing {m-n = qk + r} for some {0 \leq r < q}, we have

\displaystyle \| (m-n) \theta \|_{{\bf R}/{\bf Z}} = \| \frac{ra}{q} + (m-n) \frac{\kappa}{q} \|_{{\bf R}/{\bf Z}} \leq 2\varepsilon.

By hypothesis, {(m-n) \frac{\kappa}{q} \leq \frac{1}{10 q}}; note that as {q \leq 2/\delta} and {\varepsilon \leq 10^{-2} \delta} we also have {\varepsilon \leq \frac{1}{10q}}. This implies that {\| \frac{ra}{q} \|_{{\bf R}/{\bf Z}} < \frac{1}{q}} and thus {r=0}. We then have

\displaystyle |k \kappa| \leq 2 \varepsilon.

We conclude that for fixed {n \in I} with {\|n\theta \|_{{\bf R}/{\bf Z}} \leq \varepsilon}, there are at most {\frac{2\varepsilon}{|\kappa|}} elements {m} of {[n, n + \frac{1}{10 |\kappa|}]} such that {\|m\theta \|_{{\bf R}/{\bf Z}} \leq \varepsilon}. Iterating this with a greedy algorithm, we see that the number of {n \in I} with {\|n\theta \|_{{\bf R}/{\bf Z}} \leq \varepsilon} is at most {(\frac{N}{1/10|\kappa|} + 1) 2\varepsilon/|\kappa|}; since {\varepsilon < 10^{-2} \delta}, this implies that

\displaystyle \delta N \ll 2 \varepsilon / \kappa

and the claim follows. \Box

Now we can quickly obtain a higher degree version of Lemma 1:

Proposition 4 (Weyl exponential sum estimate, inverse form) Let {Q \subset {\bf Z}} be an arithmetic progression of length at most {N} for some {N \geq 1}, and let {P: {\bf Z} \rightarrow {\bf R}/{\bf Z}} be a polynomial of some degree at most {d \geq 0}. If

\displaystyle \frac{1}{N} |\sum_{n \in Q} e(P(n))| \geq \delta

for some {\delta > 0}, then there exists a subprogression {Q'} of {Q} with {|Q'| \gg_d \delta^{O_d(1)} N} such that {P} varies by at most {\delta} on {Q'}.

Proof: We induct on {d}. The cases {d=0,1} are immediate from Lemma 1. Now suppose that {d \geq 2}, and that the claim had already been proven for {d-1}. To simplify the notation we allow implied constants to depend on {d}. Let the hypotheses be as in the proposition. Clearly {\delta} cannot exceed {1}. By shrinking {\delta} as necessary we may assume that {\delta \leq c} for some sufficiently small constant {c} depending on {d}.

By rescaling we may assume {Q \subset [0,N] \cap {\bf Z}}. By Lemma 3, we see that for {\gg \delta^2 N} choices of {h \in [-N,N] \cap {\bf Z}} such that

\displaystyle \frac{1}{N} |\sum_{n \in I_h} e(P(n+h) - P(n))| \gg \delta^2

for some interval {I_h \subset [0,N] \cap {\bf Z}}. We write {P(n) = \sum_{i \leq d} \alpha_i n^i}, then {P(n+h)-P(n)} is a polynomial of degree at most {d-1} with leading coefficient {h \alpha_d n^{d-1}}. We conclude from induction hypothesis that for each such {h}, there exists a natural number {q_h \ll \delta^{-O(1)}} such that {\|q_h h \alpha_d \|_{{\bf R}/{\bf Z}} \ll \delta^{-O(1)} / N^{d-1}}, by double-counting, this implies that there are {\gg \delta^{O(1)} N} integers {n} in the interval {[-\delta^{-O(1)} N, \delta^{-O(1)} N] \cap {\bf Z}} such that {\|n \alpha_d \|_{{\bf R}/{\bf Z}} \ll \delta^{-O(1)} / N^{d-1}}. Applying Lemma 3, we conclude that either {N \ll \delta^{-O(1)}}, or that

\displaystyle \| q \alpha_d \|_{{\bf R}/{\bf Z}} \ll \delta^{-O(1)} / N^d. \ \ \ \ \ (4)

 

In the former case the claim is trivial (just take {Q'} to be a point), so we may assume that we are in the latter case.

We partition {Q} into arithmetic progressions {Q'} of spacing {q} and length comparable to {\delta^{-C} N} for some large {C} depending on {d} to be chosen later. By hypothesis, we have

\displaystyle \frac{1}{|Q|} |\sum_{n \in Q} e(P(n))| \geq \delta

so by the pigeonhole principle, we have

\displaystyle \frac{1}{|Q'|} |\sum_{n \in Q'} e(P(n))| \geq \delta

for at least one such progression {Q'}. On this progression, we may use the binomial theorem and (4) to write {\alpha_d n^d} as a polynomial in {n} of degree at most {d-1}, plus an error of size {O(\delta^{C - O(1)})}. We thus can write {P(n) = P'(n) + O(\delta^{C-O(1)})} for {n \in Q'} for some polynomial {P'} of degree at most {d-1}. By the triangle inequality, we thus have (for {C} large enough) that

\displaystyle \frac{1}{|Q'|} |\sum_{n \in Q'} e(P'(n))| \gg \delta

and hence by induction hypothesis we may find a subprogression {Q''} of {Q'} of size {|Q''| \gg \delta^{O(1)} N} such that {P'} varies by most {\delta/2} on {Q''}, and thus (for {C} large enough again) that {P} varies by at most {\delta} on {Q''}, and the claim follows. \Box

This gives the following corollary (also given as Exercise 16 in this previous blog post):

Corollary 5 (Weyl exponential sum estimate, inverse form II) Let {I \subset [-N,N] \cap {\bf Z}} be a discrete interval for some {N \geq 1}, and let {P(n) = \sum_{i \leq d} \alpha_i n^i} polynomial of some degree at most {d \geq 0} for some {\alpha_0,\dots,\alpha_d \in {\bf R}/{\bf Z}}. If

\displaystyle \frac{1}{N} |\sum_{n \in I} e(P(n))| \geq \delta

for some {\delta > 0}, then there is a natural number {q \ll_d \delta^{-O_d(1)}} such that {\| q\alpha_i \|_{{\bf R}/{\bf Z}} \ll_d \delta^{-O_d(1)} N^{-i}} for all {i=0,\dots,d}.

One can obtain much better exponents here using Vinogradov’s mean value theorem; see Theorem 1.6 this 2011 paper of Wooley. (Thanks to Mariusz Mirek for this reference.) However, this weaker result already suffices for many applications, and does not need any result as deep as the mean value theorem.

Proof: To simplify notation we allow implied constants to depend on {d}. As before, we may assume that {\delta \leq c} for some small constant {c>0} depending only on {d}. We may also assume that {N \geq \delta^{-C}} for some large {C}, as the claim is trivial otherwise (set {q=1}).

Applying Proposition 4, we can find a natural number {q \ll \delta^{-O(1)}} and an arithmetic subprogression {Q} of {I} such that {|Q| \gg \delta^{O(1)}} and such that {P} varies by at most {\delta} on {Q}. Writing {Q = \{ qn+r: n \in I'\}} for some interval {I' \subset [0,N] \cap {\bf Z}} of length {\gg \delta^{O(1)}} and some {0 \leq r < q}, we conclude that the polynomial {n \mapsto P(qn+r)} varies by at most {\delta} on {I'}. Taking {d^{th}} order differences, we conclude that the {d^{th}} coefficient of this polynomial is {O(\delta^{-O(1)} / N^d)}; by the binomial theorem, this implies that {n \mapsto P(qn+r)} differs by at most {O(\delta)} on {I'} from a polynomial of degree at most {d-1}. Iterating this, we conclude that the {i^{th}} coefficient of {n \mapsto P(qn+r)} is {O(\delta N^{-i})} for {i=0,\dots,d}, and the claim then follows by inverting the change of variables {n \mapsto qn+r} (and replacing {q} with a larger quantity such as {q^d} as necessary). \Box

For future reference we also record a higher degree version of the Vinogradov lemma.

Lemma 6 (Polynomial Vinogradov lemma) Let {I \subset [-N,N] \cap {\bf Z}} be a discrete interval for some {N \geq 1}, and let {P: {\bf Z} \rightarrow {\bf R}/{\bf Z}} be a polynomial {P(n) = \sum_{i \leq d} \alpha_i n^i} of degree at most {d} for some {d \geq 1} such that {\|P(n)\|_{{\bf R}/{\bf Z}} \leq \varepsilon} for at least {\delta N} values of {n \in I}, for some {0 < \varepsilon, \delta < 1}. Then either

\displaystyle N \ll_d \delta^{-O_d(1)} \ \ \ \ \ (5)

 

or

\displaystyle \varepsilon \gg_d \delta^{O_d(1)} \ \ \ \ \ (6)

 

or else there is a natural number {q \ll_d \delta^{-O_d(1)}} such that

\displaystyle \| q \alpha_i \|_{{\bf R}/{\bf Z}} \ll \frac{\delta^{-O(1)} \varepsilon}{N^i}

for all {i=0,\dots,d}.

Proof: We induct on {d}. For {d=1} this follows from Lemma 3 (noting that if {\|P(n)\|_{{\bf R}/{\bf Z}}, \|P(n_0)\|_{{\bf R}/Z} \leq \varepsilon} then {\|P(n)-P(n_0)\|_{{\bf R}/{\bf Z}} \leq 2\varepsilon}), so suppose that {d \geq 2} and that the claim is already proven for {d-1}. We now allow all implied constants to depend on {d}.

For each {h \in [-2N,2N] \cap {\bf Z}}, let {N_h} denote the number of {n \in [-N,N] \cap {\bf Z}} such that {\| P(n+h)\|_{{\bf R}/{\bf Z}}, \|P(n)\|_{{\bf R}/{\bf Z}} \leq \varepsilon}. By hypothesis, {\sum_{h \in [-2N,2N] \cap {\bf Z}} N_h \gg \delta^2 N^2}, and clearly {N_h = O(N)}, so we must have {N_h \gg \delta^2 N} for {\gg \delta^2 N} choices of {h}. For each such {h}, we then have {\|P(n+h)-P(n)\|_{{\bf R}/{\bf Z}} \leq 2\varepsilon} for {\gg \delta^2 N} choices of {n \in [-N,N] \cap {\bf Z}}, so by induction hypothesis, either (5) or (6) holds, or else for {\gg \delta^{O(1)} N} choices of {h \in [-2N,2N] \cap {\bf Z}}, there is a natural number {q_h \ll \delta^{-O(1)}} such that

\displaystyle \| q_h \alpha_{i,h} \|_{{\bf R}/{\bf Z}} \ll \frac{\delta^{-O(1)} \varepsilon}{N^i}

for {i=1,\dots,d-1}, where {\alpha_{i,h}} are the coefficients of the degree {d-1} polynomial {n \mapsto P(n+h)-P(n)}. We may of course assume it is the latter which holds. By the pigeonhole principle we may take {q_h= q} to be independent of {h}.

Since {\alpha_{d-1,h} = dh \alpha_d}, we have

\displaystyle \| qd h \alpha_d \|_{{\bf R}/{\bf Z}} \ll \frac{\delta^{-O(1)} \varepsilon}{N^{d-1}}

for {\gg \delta^{O(1)} N} choices of {h}, so by Lemma 3, either (5) or (6) holds, or else (after increasing {q} as necessary) we have

\displaystyle \| q \alpha_d \|_{{\bf R}/{\bf Z}} \ll \frac{\delta^{-O(1)} \varepsilon}{N^d}.

We can again assume it is the latter that holds. This implies that {q \alpha_{d-2,h} = (d-1) h \alpha_{d-1} + O( \delta^{-O(1)} \varepsilon / N^{d-2} )} modulo {1}, so that

\displaystyle \| q(d-1) h \alpha_{d-1} \|_{{\bf R}/{\bf Z}} \ll \frac{\delta^{-O(1)} \varepsilon}{N^{d-2}}

for {\gg \delta^{O(1)} N} choices of {h}. Arguing as before and iterating, we obtain the claim. \Box

The above results also extend to higher dimensions. Here is the higher dimensional version of Proposition 4:

Proposition 7 (Multidimensional Weyl exponential sum estimate, inverse form) Let {k \geq 1} and {N_1,\dots,N_k \geq 1}, and let {Q_i \subset {\bf Z}} be arithmetic progressions of length at most {N_i} for each {i=1,\dots,k}. Let {P: {\bf Z}^k \rightarrow {\bf R}/{\bf Z}} be a polynomial of degrees at most {d_1,\dots,d_k} in each of the {k} variables {n_1,\dots,n_k} separately. If

\displaystyle \frac{1}{N_1 \dots N_k} |\sum_{n \in Q_1 \times \dots \times Q_k} e(P(n))| \geq \delta

for some {\delta > 0}, then there exists a subprogression {Q'_i} of {Q_i} with {|Q'_i| \gg_{k,d_1,\dots,d_k} \delta^{O_{k,d_1,\dots,d_k}(1)} N_i} for each {i=1,\dots,k} such that {P} varies by at most {\delta} on {Q'_1 \times \dots \times Q'_k}.

A much more general statement, in which the polynomial phase {n \mapsto e(P(n))} is replaced by a nilsequence, and in which one does not necessarily assume the exponential sum is small, is given in Theorem 8.6 of this paper of Ben Green and myself, but it involves far more notation to even state properly.

Proof: We induct on {k}. The case {k=1} was established in Proposition 5, so we assume that {k \geq 2} and that the claim has already been proven for {k-1}. To simplify notation we allow all implied constants to depend on {k,d_1,\dots,d_k}. We may assume that {\delta \leq c} for some small {c>0} depending only on {k,d_1,\dots,d_k}.

By a linear change of variables, we may assume that {Q_i \subset [0,N_i] \cap {\bf Z}} for all {i=1,\dots,k}.

We write {n' := (n_1,\dots,n_{k-1})}. First suppose that {N_k = O(\delta^{-O(1)})}. Then by the pigeonhole principle we can find {n_k \in I_k} such that

\displaystyle \frac{1}{N_1 \dots N_{k-1}} |\sum_{n' \in Q_1 \times \dots \times Q_{k-1}} e(P(n',n_k))| \geq \delta

and the claim then follows from the induction hypothesis. Thus we may assume that {N_k \geq \delta^{-C}} for some large {C} depending only on {k,d_1,\dots,d_k}. Similarly we may assume that {N_i \geq \delta^{-C}} for all {i=1,\dots,k}.

By the triangle inequality, we have

\displaystyle \frac{1}{N_1 \dots N_k} \sum_{n_k \in Q_k} |\sum_{n' \in Q_1 \times \dots \times Q_{k-1}} e(P(n',n_k))| \geq \delta.

The inner sum is {O(N_k)}, and the outer sum has {O(N_1 \dots N_{k-1})} terms. Thus, for {\gg \delta N_1 \dots N_{k-1}} choices of {n' \in Q_1 \times \dots \times Q_{k-1}}, one has

\displaystyle \frac{1}{N_k} |\sum_{n_k \in Q_k} e(P(n',n_k))| \gg \delta. \ \ \ \ \ (7)

 

We write

\displaystyle P(n',n_k) = \sum_{i_k \leq d_k} P_{i_k}(n') n_k^i

for some polynomials {P_{i_k}: {\bf Z}^{k-1} \rightarrow {\bf R}/{\bf Z}} of degrees at most {d_1,\dots,d_{k-1}} in the variables {n_1,\dots,n_{k-1}}. For each {n'} obeying (7), we apply Corollary 5 to conclude that there exists a natural number {q_{n'} \ll \delta^{-O(1)}} such that

\displaystyle \| q_{n'} P_{i_k}(n') \|_{{\bf R}/{\bf Z}} \ll \delta^{-O(1)} / N_k^{i_k}

for {i_k=1,\dots,d_k} (the claim also holds for {i_k=0} but we discard it as being trivial). By the pigeonhole principle, there thus exists a natural number {q \ll \delta^{-O(1)}} such that

\displaystyle \| q P_{i_k}(n') \|_{{\bf R}/{\bf Z}} \ll \delta^{-O(1)} / N_k^{i_k}

for all {i_k=1,\dots,d_k} and for {\gg \delta^{O(1)} N_1 \dots N_{k-1}} choices of {n' \in Q_1 \times \dots \times Q_{k-1}}. If we write

\displaystyle P_{i_k}(n') = \sum_{i_{k-1} \leq d_{k-1}} P_{i_{k-1},i_k}(n_1,\dots,n_{k-2}) n_{k-1}^{i_{k-1}},

where {P_{i_{k-1},i_k}: {\bf Z}^{k-2} \rightarrow {\bf R}/{\bf Z}} is a polynomial of degrees at most {d_1,\dots,d_{k-2}}, then for {\gg \delta^{O(1)} N_1 \dots N_{k-2}} choices of {(n_1,\dots,n_{k-2}) \in Q_1 \times \dots \times Q_{k-2}} we then have

\displaystyle \| \sum_{i_{k-1} \leq d_{k-1}} q P_{i_{k-1},i_k}(n_1,\dots,n_{k-2}) n_{k-1}^{i_{k-1}} \|_{{\bf R}/{\bf Z}} \ll \delta^{-O(1)} / N_k^{i_k}.

Applying Lemma 6 in the {n_{k-1}} and the largeness hypotheses on the {N_i} (and also the assumption that {i_k \geq 1}) we conclude (after enlarging {q} as necessary, and pigeonholing to keep {q} independent of {n_1,\dots,n_{k-2}}) that

\displaystyle \| q P_{i_{k-1},i_k}(n_1,\dots,n_{k-2}) \|_{{\bf R}/{\bf Z}} \ll \frac{\delta^{-O(1)}}{N_{k-1}^{i_{k-1}} N_k^{i_k}}

for all {i_{k-1}=0,\dots,d_{k-1}} (note that we now include that {i_{k-1}=0} case, which is no longer trivial) and for {\gg \delta^{O(1)} N_1 \dots N_{k-2}} choices of {(n_1,\dots,n_{k-2}) \in Q_1 \times \dots \times Q_{k-2}}. Iterating this, we eventually conclude (after enlarging {q} as necessary) that

\displaystyle \| q \alpha_{i_1,\dots,i_k} \|_{{\bf R}/{\bf Z}} \ll \frac{\delta^{-O(1)}}{N_1^{i_1} \dots N_k^{i_k}} \ \ \ \ \ (8)

 

whenever {i_j \in \{0,\dots,d_j\}} for {j=1,\dots,k}, with {i_k} nonzero. Permuting the indices, and observing that the claim is trivial for {(i_1,\dots,i_k) = (0,\dots,0)}, we in fact obtain (8) for all {(i_1,\dots,i_k) \in \{0,\dots,d_1\} \times \dots \times \{0,\dots,d_k\}}, at which point the claim easily follows by taking {Q'_j := \{ qn_j: n_j \leq \delta^C N_j\}} for each {j=1,\dots,k}. \Box

An inspection of the proof of the above result (or alternatively, by combining the above result again with many applications of Lemma 6) reveals the following general form of Proposition 4, which was posed as Exercise 17 in this previous blog post, but had a slight misprint in it (it did not properly treat the possibility that some of the {N_j} could be small) and was a bit trickier to prove than anticipated (in fact, the reason for this post was that I was asked to supply a more detailed solution for this exercise):

Proposition 8 (Multidimensional Weyl exponential sum estimate, inverse form, II) Let {k \geq 1} be an natural number, and for each {j=1,\dots,k}, let {I_j \subset [0,N_j]_{\bf Z}} be a discrete interval for some {N_j \geq 1}. Let

\displaystyle P(n_1,\dots,n_k) = \sum_{i_1 \leq d_1, \dots, i_k \leq d_k} \alpha_{i_1,\dots,i_k} n_1^{i_1} \dots n_k^{i_k}

be a polynomial in {k} variables of multidegrees {d_1,\dots,d_k \geq 0} for some {\alpha_{i_1,\dots,i_k} \in {\bf R}/{\bf Z}}. If

\displaystyle \frac{1}{N_1 \dots N_k} |\sum_{n \in I_1 \times \dots \times I_k} e(P(n))| \geq \delta \ \ \ \ \ (9)

 

for some {\delta > 0}, then either

\displaystyle N_j \ll_{k,d_1,\dots,d_k} \delta^{-O_{k,d_1,\dots,d_k}(1)} \ \ \ \ \ (10)

 

for some {1 \leq j \leq d_k}, or else there is a natural number {q \ll_{k,d_1,\dots,d_k} \delta^{-O_{k,d_1,\dots,d_k}(1)}} such that

\displaystyle \| q\alpha_{i_1,\dots,i_k} \|_{{\bf R}/{\bf Z}} \ll_{k,d_1,\dots,d_k} \delta^{-O_d(1)} N_1^{-i_1} \dots N_k^{-i_k} \ \ \ \ \ (11)

 

whenever {i_j \leq d_j} for {j=1,\dots,k}.

Again, the factor of {N_1^{-i_1} \dots N_k^{-i_k}} is natural in this bound. In the {k=1} case, the option (10) may be deleted since (11) trivially holds in this case, but this simplification is no longer available for {k>1} since one needs (10) to hold for all {j} (not just one {j}) to make (11) completely trivial. Indeed, the above proposition fails for {k \geq 2} if one removes (10) completely, as can be seen for instance by inspecting the exponential sum {\sum_{n_1 \in \{0,1\}} \sum_{n_2 \in [1,N] \cap {\bf Z}} e( \alpha n_1 n_2)}, which has size comparable to {N} regardless of how irrational {\alpha} is.

Vitaly Bergelson, Tamar Ziegler, and I have just uploaded to the arXiv our joint paper “Multiple recurrence and convergence results associated to {{\bf F}_{p}^{\omega}}-actions“. This paper is primarily concerned with limit formulae in the theory of multiple recurrence in ergodic theory. Perhaps the most basic formula of this type is the mean ergodic theorem, which (among other things) asserts that if {(X,{\mathcal X}, \mu,T)} is a measure-preserving {{\bf Z}}-system (which, in this post, means that {(X,{\mathcal X}, \mu)} is a probability space and {T: X \mapsto X} is measure-preserving and invertible, thus giving an action {(T^n)_{n \in {\bf Z}}} of the integers), and {f,g \in L^2(X,{\mathcal X}, \mu)} are functions, and {X} is ergodic (which means that {L^2(X,{\mathcal X}, \mu)} contains no {T}-invariant functions other than the constants (up to almost everywhere equivalence, of course)), then the average

\displaystyle  \frac{1}{N} \sum_{n=1}^N \int_X f(x) g(T^n x)\ d\mu \ \ \ \ \ (1)

converges as {N \rightarrow \infty} to the expression

\displaystyle  (\int_X f(x)\ d\mu) (\int_X g(x)\ d\mu);

see e.g. this previous blog post. Informally, one can interpret this limit formula as an equidistribution result: if {x} is drawn at random from {X} (using the probability measure {\mu}), and {n} is drawn at random from {\{1,\ldots,N\}} for some large {N}, then the pair {(x, T^n x)} becomes uniformly distributed in the product space {X \times X} (using product measure {\mu \times \mu}) in the limit as {N \rightarrow \infty}.

If we allow {(X,\mu)} to be non-ergodic, then we still have a limit formula, but it is a bit more complicated. Let {{\mathcal X}^T} be the {T}-invariant measurable sets in {{\mathcal X}}; the {{\bf Z}}-system {(X, {\mathcal X}^T, \mu, T)} can then be viewed as a factor of the original system {(X, {\mathcal X}, \mu, T)}, which is equivalent (in the sense of measure-preserving systems) to a trivial system {(Z_0, {\mathcal Z}_0, \mu_{Z_0}, 1)} (known as the invariant factor) in which the shift is trivial. There is then a projection map {\pi_0: X \rightarrow Z_0} to the invariant factor which is a factor map, and the average (1) converges in the limit to the expression

\displaystyle  \int_{Z_0} (\pi_0)_* f(z) (\pi_0)_* g(z)\ d\mu_{Z_0}(x), \ \ \ \ \ (2)

where {(\pi_0)_*: L^2(X,{\mathcal X},\mu) \rightarrow L^2(Z_0,{\mathcal Z}_0,\mu_{Z_0})} is the pushforward map associated to the map {\pi_0: X \rightarrow Z_0}; see e.g. this previous blog post. We can interpret this as an equidistribution result. If {(x,T^n x)} is a pair as before, then we no longer expect complete equidistribution in {X \times X} in the non-ergodic, because there are now non-trivial constraints relating {x} with {T^n x}; indeed, for any {T}-invariant function {f: X \rightarrow {\bf C}}, we have the constraint {f(x) = f(T^n x)}; putting all these constraints together we see that {\pi_0(x) = \pi_0(T^n x)} (for almost every {x}, at least). The limit (2) can be viewed as an assertion that this constraint {\pi_0(x) = \pi_0(T^n x)} are in some sense the “only” constraints between {x} and {T^n x}, and that the pair {(x,T^n x)} is uniformly distributed relative to these constraints.

Limit formulae are known for multiple ergodic averages as well, although the statement becomes more complicated. For instance, consider the expression

\displaystyle  \frac{1}{N} \sum_{n=1}^N \int_X f(x) g(T^n x) h(T^{2n} x)\ d\mu \ \ \ \ \ (3)

for three functions {f,g,h \in L^\infty(X, {\mathcal X}, \mu)}; this is analogous to the combinatorial task of counting length three progressions in various sets. For simplicity we assume the system {(X,{\mathcal X},\mu,T)} to be ergodic. Naively one might expect this limit to then converge to

\displaystyle  (\int_X f\ d\mu) (\int_X g\ d\mu) (\int_X h\ d\mu)

which would roughly speaking correspond to an assertion that the triplet {(x,T^n x, T^{2n} x)} is asymptotically equidistributed in {X \times X \times X}. However, even in the ergodic case there can be additional constraints on this triplet that cannot be seen at the level of the individual pairs {(x,T^n x)}, {(x, T^{2n} x)}. The key obstruction here is that of eigenfunctions of the shift {T: X \rightarrow X}, that is to say non-trivial functions {f: X \rightarrow S^1} that obey the eigenfunction equation {Tf = \lambda f} almost everywhere for some constant (or {T}-invariant) {\lambda}. Each such eigenfunction generates a constraint

\displaystyle  f(x) \overline{f(T^n x)}^2 f(T^{2n} x) = 1 \ \ \ \ \ (4)

tying together {x}, {T^n x}, and {T^{2n} x}. However, it turns out that these are in some sense the only constraints on {x,T^n x, T^{2n} x} that are relevant for the limit (3). More precisely, if one sets {{\mathcal X}_1} to be the sub-algebra of {{\mathcal X}} generated by the eigenfunctions of {T}, then it turns out that the factor {(X, {\mathcal X}_1, \mu, T)} is isomorphic to a shift system {(Z_1, {\mathcal Z}_1, \mu_{Z_1}, x \mapsto x+\alpha)} known as the Kronecker factor, for some compact abelian group {Z_1 = (Z_1,+)} and some (irrational) shift {\alpha \in Z_1}; the factor map {\pi_1: X \rightarrow Z_1} pushes eigenfunctions forward to (affine) characters on {Z_1}. It is then known that the limit of (3) is

\displaystyle  \int_\Sigma (\pi_1)_* f(x_0) (\pi_1)_* g(x_1) (\pi_1)_* h(x_2)\ d\mu_\Sigma

where {\Sigma \subset Z_1^3} is the closed subgroup

\displaystyle  \Sigma = \{ (x_1,x_2,x_3) \in Z_1^3: x_1-2x_2+x_3=0 \}

and {\mu_\Sigma} is the Haar probability measure on {\Sigma}; see this previous blog post. The equation {x_1-2x_2+x_3=0} defining {\Sigma} corresponds to the constraint (4) mentioned earlier. Among other things, this limit formula implies Roth’s theorem, which in the context of ergodic theory is the assertion that the limit (or at least the limit inferior) of (3) is positive when {f=g=h} is non-negative and not identically vanishing.

If one considers a quadruple average

\displaystyle  \frac{1}{N} \sum_{n=1}^N \int_X f(x) g(T^n x) h(T^{2n} x) k(T^{3n} x)\ d\mu \ \ \ \ \ (5)

(analogous to counting length four progressions) then the situation becomes more complicated still, even in the ergodic case. In addition to the (linear) eigenfunctions that already showed up in the computation of the triple average (3), a new type of constraint also arises from quadratic eigenfunctions {f: X \rightarrow S^1}, which obey an eigenfunction equation {Tf = \lambda f} in which {\lambda} is no longer constant, but is now a linear eigenfunction. For such functions, {f(T^n x)} behaves quadratically in {n}, and one can compute the existence of a constraint

\displaystyle  f(x) \overline{f(T^n x)}^3 f(T^{2n} x)^3 \overline{f(T^{3n} x)} = 1 \ \ \ \ \ (6)

between {x}, {T^n x}, {T^{2n} x}, and {T^{3n} x} that is not detected at the triple average level. As it turns out, this is not the only type of constraint relevant for (5); there is a more general class of constraint involving two-step nilsystems which we will not detail here, but see e.g. this previous blog post for more discussion. Nevertheless there is still a similar limit formula to previous examples, involving a special factor {(Z_2, {\mathcal Z}_2, \mu_{Z_2}, S)} which turns out to be an inverse limit of two-step nilsystems; this limit theorem can be extracted from the structural theory in this paper of Host and Kra combined with a limit formula for nilsystems obtained by Lesigne, but will not be reproduced here. The pattern continues to higher averages (and higher step nilsystems); this was first done explicitly by Ziegler, and can also in principle be extracted from the structural theory of Host-Kra combined with nilsystem equidistribution results of Leibman. These sorts of limit formulae can lead to various recurrence results refining Roth’s theorem in various ways; see this paper of Bergelson, Host, and Kra for some examples of this.

The above discussion was concerned with {{\bf Z}}-systems, but one can adapt much of the theory to measure-preserving {G}-systems for other discrete countable abelian groups {G}, in which one now has a family {(T_g)_{g \in G}} of shifts indexed by {G} rather than a single shift, obeying the compatibility relation {T_{g+h}=T_g T_h}. The role of the intervals {\{1,\ldots,N\}} in this more general setting is replaced by that of Folner sequences. For arbitrary countable abelian {G}, the theory for double averages (1) and triple limits (3) is essentially identical to the {{\bf Z}}-system case. But when one turns to quadruple and higher limits, the situation becomes more complicated (and, for arbitrary {G}, still not fully understood). However one model case which is now well understood is the finite field case when {G = {\bf F}_p^\omega = \bigcup_{n=1}^\infty {\bf F}_p^n} is an infinite-dimensional vector space over a finite field {{\bf F}_p} (with the finite subspaces {{\bf F}_p^n} then being a good choice for the Folner sequence). Here, the analogue of the structural theory of Host and Kra was worked out by Vitaly, Tamar, and myself in these previous papers (treating the high characteristic and low characteristic cases respectively). In the finite field setting, it turns out that nilsystems no longer appear, and one only needs to deal with linear, quadratic, and higher order eigenfunctions (known collectively as phase polynomials). It is then natural to look for a limit formula that asserts, roughly speaking, that if {x} is drawn at random from a {{\bf F}_p^\omega}-system and {n} drawn randomly from a large subspace of {{\bf F}_p^\omega}, then the only constraints between {x, T^n x, \ldots, T^{(p-1)n} x} are those that arise from phase polynomials. The main theorem of this paper is to establish this limit formula (which, again, is a little complicated to state explicitly and will not be done here). In particular, we establish for the first time that the limit actually exists (a result which, for {{\bf Z}}-systems, was one of the main results of this paper of Host and Kra).

As a consequence, we can recover finite field analogues of most of the results of Bergelson-Host-Kra, though interestingly some of the counterexamples demonstrating sharpness of their results for {{\bf Z}}-systems (based on Behrend set constructions) do not seem to be present in the finite field setting (cf. this previous blog post on the cap set problem). In particular, we are able to largely settle the question of when one has a Khintchine-type theorem that asserts that for any measurable set {A} in an ergodic {{\bf F}_p^\omega}-system and any {\epsilon>0}, one has

\displaystyle  \mu( T_{c_1 n} A \cap \ldots \cap T_{c_k n} A ) > \mu(A)^k - \epsilon

for a syndetic set of {n}, where {c_1,\ldots,c_k \in {\bf F}_p} are distinct residue classes. It turns out that Khintchine-type theorems always hold for {k=1,2,3} (and for {k=1,2} ergodicity is not required), and for {k=4} it holds whenever {c_1,c_2,c_3,c_4} form a parallelogram, but not otherwise (though the counterexample here was such a painful computation that we ended up removing it from the paper, and may end up putting it online somewhere instead), and for larger {k} we could show that the Khintchine property failed for generic choices of {c_1,\ldots,c_k}, though the problem of determining exactly the tuples for which the Khintchine property failed looked to be rather messy and we did not completely settle it.

Archives