You are currently browsing the monthly archive for December 2024.
Hamilton’s quaternion number system is a non-commutative extension of the complex numbers, consisting of numbers of the form
where
are real numbers, and
are anti-commuting square roots of
with
,
,
. While they are non-commutative, they do keep many other properties of the complex numbers:
- Being non-commutative, the quaternions do not form a field. However, they are still a skew field (or division ring): multiplication is associative, and every non-zero quaternion has a unique multiplicative inverse.
- Like the complex numbers, the quaternions have a conjugation
although this is now an antihomomorphism rather than a homomorphism:. One can then split up a quaternion
into its real part
and imaginary part
by the familiar formulae
(though we now leave the imaginary part purely imaginary, as opposed to dividing byin the complex case).
- The inner product
is symmetric and positive definite (withforming an orthonormal basis). Also, for any
,
is real, hence equal to
. Thus we have a norm
Since the real numbers commute with all quaternions, we have the multiplicative property. In particular, the unit quaternions
(also known as
,
, or
) form a compact group.
- We have the cyclic trace property
which allows one to take adjoints of left and right multiplication: - As
are square roots of
, we have the usual Euler formulae
for real, together with other familiar formulae such as
,
,
, etc.
The unit quaternions act on the imaginary quaternions
by conjugation:
For instance, for any real , conjugation by
is a rotation by
around
:
. The doubling of the angle here can be explained from the Lie algebra fact that
is
rather than
; it also closely related to the aforementioned double cover. We also of course have
acting on
by left multiplication; this is known as the spinor representation, but will not be utilized much in this post. (Giving
the right action of
makes it a copy of
, and the spinor representation then also becomes the standard representation of
on
.)
Given how quaternions relate to three-dimensional rotations, it is not surprising that one can also be used to recover the basic laws of spherical trigonometry – the study of spherical triangles on the unit sphere. This is fairly well known, but it took a little effort for me to locate the required arguments, so I am recording the calculations here.
The first observation is that every unit quaternion induces a unit tangent vector
on the unit sphere
, located at
; the third unit vector
is then another tangent vector orthogonal to the first two (and oriented to the left of the original tangent vector), and can be viewed as the cross product of
and
. Right multplication of this quaternion then corresponds to various natural operations on this unit tangent vector:
- Right multiplying
by
does not affect the location
of the tangent vector, but rotates the tangent vector
anticlockwise by
in the direction of the orthogonal tangent vector
, as it replaces
by
.
- Right multiplying
by
advances the tangent vector by geodesic flow by angle
, as it replaces
by
, and replaces
by
.
Now suppose one has a spherical triangle with vertices , with the spherical arcs
subtending angles
respectively, and the vertices
subtending angles
respectively; suppose also that
is oriented in an anti-clockwise direction for sake of discussion. Observe that if one starts at
with a tangent vector oriented towards
, advances that vector by
, and then rotates by
, the tangent vector now at
and pointing towards
. If one advances by
and rotates by
, one is now at
pointing towards
; and if one then advances by
and rotates by
, one is back at
pointing towards
. This gives the fundamental relation
action, the right-hand side could conceivably have been
rather than
; but for extremely small triangles the right-hand side is clearly
, and so by continuity it must be
for all triangles.) Indeed, a moments thought will reveal that the condition (4) is necessary and sufficient for the data
to be associated with a spherical triangle. Thus one can view (4) as a “master equation” for spherical trigonometry: in principle, it can be used to derive all the other laws of this subject.
Remark 1 The law (4) has an evident symmetry, which corresponds to the operation of replacing a spherical triangle with its dual triangle. Also, there is nothing particularly special about the choice of imaginaries
in (4); one can conjugate (4) by various quaternions and replace
here by any other orthogonal pair of unit quaternions.
Remark 2 If we work in the small scale regime, replacingby
for some small
, then we expect spherical triangles to behave like Euclidean triangles. Indeed, (4) to zeroth order becomes
which reflects the classical fact that the sum of angles of a Euclidean triangle is equal to
. To first order, one obtains
which reflects the evident fact that the vector sum of the sides of a Euclidean triangle sum to zero. (Geometrically, this correspondence reflects the fact that the action of the (projective) quaternion group on the unit sphere converges to the action of the special Euclidean group
on the plane, in a suitable asymptotic limit.)
The identity (4) is an identity of two unit quaternions; as the unit quaternion group is three-dimensional, this thus imposes three independent constraints on the six real parameters
of the spherical triangle. One can manipulate this constraint in various ways to obtain various trigonometric identities involving some subsets of these six parameters. For instance, one can rearrange (4) to get
to reverse the sign of
, we also have
In a similar fashion, from (5) we see that the quantity
As a variant of the above analysis, we have from (5) again that
Example 3 One application of Napier’s rule (6) is to determine the sunrise equation for when the sun rises and sets at a given location on the Earth, and a given time of year. For sake of argument let us work in summer, in which the declinationof the Sun is positive (due to axial tilt, it reaches a maximum of
at the summer solstice). Then the Sun subtends an angle of
from the pole star (Polaris in the northern hemisphere, Sigma Octantis in the southern hemisphere), and appears to rotate around that pole star once every
hours. On the other hand, if one is at a latitude
, then the pole star an elevation of
above the horizon. At extremely high latitudes
, the sun will never set (a phenomenon known as “midnight sun“); but in all other cases, at sunrise or sunset, the sun, pole star, and horizon point below the pole star will form a right-angled spherical triangle, with hypotenuse subtending an angle
and vertical side subtending an angle
. The angle subtended by the pole star in this triangle is
, where
is the solar hour angle
– the angle that the sun deviates from its noon position. Equation (6) then gives the sunrise equation
or equivalently
A similar rule determines the time of sunset. In particular, the number of daylight hours in summer (assuming one is not in the midnight sun scenario
) is given by
The situation in winter is similar, except that
is now negative, and polar night (no sunrise) occurs when
.
I’ve just uploaded to the arXiv the paper “On the distribution of eigenvalues of GUE and its minors at fixed index“. This is a somewhat technical paper establishing some estimates regarding one of the most well-studied random matrix models, the Gaussian Unitary Ensemble (GUE), that were not previously in the literature, but which will be needed for some forthcoming work of Hariharan Narayanan on the limiting behavior of “hives” with GUE boundary conditions (building upon our previous joint work with Sheffield).
For sake of discussion we normalize the GUE model to be the random Hermitian matrix
whose probability density function is proportional to
. With this normalization, the famous Wigner semicircle law will tell us that the eigenvalues
of this matrix will almost all lie in the interval
, and after dividing by
, will asymptotically be distributed according to the semicircle distribution
Eigenvalues can be described by their index or by their (normalized) energy
. In principle, the two descriptions are related by the classical map
defined above, but there are microscopic fluctuations from the classical location that create subtle technical difficulties between “fixed index” results in which one focuses on a single index
(and neighboring indices
, etc.), and “fixed energy” results in which one focuses on a single energy
(and eigenvalues near this energy). The phenomenon of eigenvalue rigidity does give some control on these fluctuations, allowing one to relate “averaged index” results (in which the index
ranges over a mesoscopic range) with “averaged energy” results (in which the energy
is similarly averaged over a mesoscopic interval), but there are technical issues in passing back from averaged control to pointwise control, either for the index or energy.
We will be mostly concerned in the bulk region where the index is in an inteval of the form
for some fixed
, or equivalently the energy
is in
for some fixed
. In this region it is natural to introduce the normalized eigenvalue gaps
However, these results left open the possibility of bad tail behavior at extremely large or small values of the gaps ; in particular, moments of the
were not directly controlled by previous results. The first result of the paper is to push the determinantal analysis further, and obtain such results. For instance, we obtain moment bounds
A key point in these estimates is that no factors of occur in the estimates, which is what one would obtain if one tried to use existing eigenvalue rigidity theorems. (In particular, if one normalized the eigenvalues
at the same scale at the gap
, they would fluctuate by a standard deviation of about
; it is only the gaps between eigenvalues that exhibit much smaller fluctuation.) On the other hand, the dependence on
is not optimal, although it was sufficient for the applications I had in mind.
As with my previous paper, the strategy is to try to replace fixed index events such as with averaged energy events. For instance, if
and
has classical location
, then there is an interval of normalized energies
of length
, with the property that there are precisely
eigenvalues to the right of
and no eigenvalues in the interval
, where
For the intended application to GUE hives, it is important to not just control gaps of the eigenvalues
of the GUE matrix
, but also the gaps
of the eigenvalues
of the top left
minor
of
. This minor of a GUE matrix is basically again a GUE matrix, so the above theorem applies verbatim to the
; but it turns out to be necessary to control the joint distribution of the
and
, and also of the interlacing gaps
between the
and
. For fixed energy, these gaps are in principle well understood, due to previous work of Adler-Nordenstam-van Moerbeke and of Johansson-Nordenstam which show that the spectrum of both matrices is asymptotically controlled by the Boutillier bead process. This also gives averaged energy and averaged index results without much difficulty, but to get to fixed index information, one needs some universality result in the index
. For the gaps
of the original matrix, such a universality result is available due to the aforementioned work of Erdos and Yau, but this does not immediately imply the corresponding universality result for the joint distribution of
and
or
. For this, we need a way to relate the eigenvalues
of the matrix
to the eigenvalues
of the minors
. By a standard Schur’s complement calculation, one can obtain the equation
This at last brings us to the final result of the paper, which is the one which is actually needed for the application to GUE hives. Here, one is interested in controlling the variance of a linear combination of a fixed number
of consecutive interlacing gaps
, where the
are arbitrary deterministic coefficients. An application of the triangle and Cauchy-Schwarz inequalities, combined with the previous moment bounds on gaps, shows that this randomv ariable has variance
. However, this bound is not expected to be sharp, due to the expected decay between correlations of eigenvalue gaps. In this paper, I improve the variance bound to
This improvement reflects some decay in the covariances between distant interlacing gaps . I was not able to establish such decay directly. Instead, using some Fourier analysis, one can reduce matters to studying the case of modulated linear statistics such as
for various frequencies
. In “high frequency” cases one can use the triangle inequality to reduce matters to studying the original eigenvalue gaps
, which can be handled by a (somewhat complicated) determinantal process calculation, after first using universality results to pass from fixed index to averaged index, thence to averaged energy, then to fixed energy estimates. For low frequencies the triangle inequality argument is unfavorable, and one has to instead use the determinantal kernel of the full minor process, and not just an individual matrix. This requires some classical, but tedious, calculation of certain asymptotics of sums involving Hermite polynomials.
The full argument is unfortunately quite complex, but it seems that the combination of having to deal with minors, as well as fixed indices, places this result out of reach of many prior methods.
Renaissance Philanthropy and XTX Markets have announced the launch of the AI for Math Fund, a new grant program supporting projects that apply AI and machine learning to mathematics, with a focus on automated theorem proving, with an initial $9.2 million in funding. The project funding categories, and examples of projects in such categories, are:
1. Production Grade Software Tools
- AI-based autoformalization tools for translating natural-language mathematics into the formalisms of proof assistants
- AI-based auto-informalization tools for translating proof-assistant proofs into interpretable natural-language mathematics
- AI-based models for suggesting tactics/steps or relevant concepts to the user of a proof assistant, or for generating entire proofs
- Infrastructure to connect proof assistants with computer algebra systems, calculus, and PDEs
- A large-scale, AI-enhanced distributed collaboration platform for mathematicians
2. Datasets
- Datasets of formalized theorems and proofs in a proof assistant
- Datasets that would advance AI for theorem proving as applied to program verification and secure code generation
- Datasets of (natural-language) mathematical problems, theorems, proofs, exposition, etc.
- Benchmarks and training environments associated with datasets and model tasks (autoformalization, premise selection, tactic or proof generation, etc.)
3. Field Building
- Textbooks
- Courses
- Documentation and support for proof assistants, and interfaces/APIs to integrate with AI tools
4. Breakthrough Ideas
- Expected difficulty estimation (of sub-problems of a proof)
- Novel mathematical implications of proofs formalized type-theoretically
- Formalization of proof complexity in proof assistants
The deadline for initial expressions of interest is Jan 10, 2025.
[Disclosure: I have agreed to serve on the advisory board for this fund.]
Update: See also this discussion thread on possible projects that might be supported by this fund.


Recent Comments