Introduction
The main results
The long-standing twin prime conjecture asserts that, if \(p_n\) is the \(n\)-th prime, then \[\liminf_{n\to\infty}(p_{n+1}-p_n)=2.\] One of many reasons for the difficulty of the conjecture is the fact that it imposes additive structure on primes, which are defined multiplicatively. The twin prime conjecture is a special case of a more general conjecture, the Hardy–Littlewood \(k\)-tuples conjecture, which we now state. Suppose that \(\mathcal{H}_k = \{h_1,\ldots,h_k\}\) is an admissible set (Definition 1). The conjecture then states that there are infinitely many positive integers \(n\) such that each of the numbers \(n+h_1,\ldots,n+h_k\) is prime. Since \(\{0,2\}\) is an admissible set, the Hardy–Littlewood \(k\)-tuples conjecture implies the twin prime conjecture.
The prime number theorem implies that as \(n\to\infty\), the average value of \(p_{n+1}-p_n\) is asymptotic to \(\log p_n\). Foundational work of Goldston, Pintz, and Yıldırım developed the “GPY method,” which led to the proof that \[\liminf_{n\to\infty}\frac{p_{n+1}-p_n}{\log p_n}=0.\] Their method crucially relies on the distribution of primes in arithmetic progressions given by the Bombieri–Vinogradov theorem: If \(0<\varepsilon<\frac{1}{2}\) and \(A>0\) are fixed, then for all \(x\geq 2\), \[\sum_{q\leq x^{1/2-\varepsilon}}\max_{\gcd(a,q)=1}\Big|\pi(x;q,a)-\frac{\pi(x)}{\varphi(q)}\Big|\ll_A \frac{x}{(\log x)^A}.\] The Elliott–Halberstam conjecture asserts that \(x^{1/2-\varepsilon}\) can be replaced with \(x^{1-\varepsilon}\). If there exists a fixed \(\delta>0\) such that \(1/2-\varepsilon\) can be improved to \(1/2+\delta\), then the GPY method suffices to establish the infinitude of bounded gaps between primes, namely \[\liminf_{n\to\infty}(p_{n+1}-p_n)<\infty.\]
In May 2013, Zhang established deep new results on the distribution of primes in arithmetic progressions, a variant of the Bombieri–Vinogradov theorem. When combined with a modification of the GPY method, it yields for the first time the existence of infinitely many bounded gaps between primes: \[\liminf_{n\to\infty}(p_{n+1}-p_n)\leq 70\,000\,000.\] The work of Zhang was meticulously refined by the Polymath 8a project, the net result of which was the bound \[\liminf_{n\to\infty}(p_{n+1}-p_n)\leq 4\,680.\]
In November 2013, the breakthrough work of Maynard1 established a substantially more robust version of the GPY method that leads to significant theoretical and numerical improvements over the work of Zhang using only the original Bombieri–Vinogradov theorem. Writing \[H_m:=\liminf_{n\to\infty}(p_{m+n}-p_n),\] Maynard proved that \(H_m\ll m^3 e^{4m}\), and in particular that \(H_1\leq 600\). Shortly thereafter, the Polymath8b project theoretically and numerically refined these ideas, obtaining \(H_m \ll m e^{(4-\frac{28}{157})m}\) and in particular \(H_1\leq 246\).
This blueprint develops both explicit bounds from a single multidimensional sieve. The only result that is assumed is the Bombieri–Vinogradov theorem.
Theorem 1 (First main result). Assume the Bombieri–Vinogradov theorem. Then \(H_1 \le 600\); that is, \[\liminf_{n\to\infty}(p_{n+1}-p_n)\le 600.\]
Theorem 2 (Second main result). Assume the Bombieri–Vinogradov theorem. Then \(H_1 \le 246\); that is, \[\liminf_{n\to\infty}(p_{n+1}-p_n)\le 246.\]
Blueprinting bounded gaps
This blueprint is designed to facilitate a formalisation of the
bounds \(H_1\leq 600\) and \(H_1\leq 246\). It exhibits several key
differences from a typical research article, some of which we now
outline. For convenience, we call lemmas, sublemmas, propositions, and
theorems provable statements; a definition
or notation is a declaration; and an
entity is a provable statement or a
declaration. Given an entity, all other entities that are referenced in
the statement (resp. proof) of the entity are catalogued in a
uses-notation (resp. uses-proof) environment.
A uses statement is a
uses-notation or uses-proof environment.
Each declaration defines exactly one object. Within each declaration environment, only one declaration is introduced. No declaration environment contains any provable statement. Conversely, a provable statement only contains statements to be proved; no declarations are introduced in any provable statement. When recalling a declaration, only declarations are referenced, not provable statements.
Each provable statement has a minimal set of hypotheses under which exactly one stated conclusion holds. Multiple conclusions that might be bundled together must be separated into smaller provable statements.
Provable statements are decomposed into smaller, “atomic” provable statements, each of which will ideally correspond with a single formal declaration.
The dependency graph of our proof is directed and absent of any cycles. The edges of this graph are indicated by our uses statements. An entity may only cite entities that come before; no “forward-referencing” is permitted. A uses statement associated to a declaration cannot cite a provable statement.
Apart from the main target theorems, every entity must be cited by at least one uses statement.
Every reference inside a provable statement or its proof must appear in the uses statement of said provable statement.
Superfluous or duplicated entities are deleted. In particular, there are no
remarkorcorollaryenvironments.Imported or external results are stated once in a dedicated “external inputs” section and explicitly cited at least once in a uses statement.
The exposition is as “linear” as feasible, which we loosely describe as follows.
Declarations come first.
Lemmas and sublemmas come next; they incorporate the declarations that came before. Lemmas are not proved before the entities in their uses statements are given (and proved, if necessary).
Propositions are then proved using the lemmas and declarations. Propositions are not proved before the entities in their uses statements are given (and proved, if necessary).
The main theorems are then proved using all pertinent entities that precede them. A main theorem is not proved before the entities in its uses statement are given (and proved, if necessary).
Overview
We give a blueprint of the proofs of the two bounds of Section 1, in a form meant to be formalized in Lean. This section outlines the method; the rest of the paper carries it out. Throughout, the only analytic input is the Bombieri–Vinogradov theorem.
The GPY method and the sum \(S_2(\lambda) - \rho S_1(\lambda)\)
Fix an admissible \(k\)-tuple \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) (Definition 1); we seek infinitely many \(n\) for which at least \(\lfloor\rho\rfloor + 1\) of \(n+h_1, \dotsc, n+h_k\) are prime. Following Goldston–Pintz–Yıldırım, introduce weights \(w_n \ge 0\) and set \[S(N, \rho) \;:=\; \sum_{N < n \le 2N}\Bigl(\sum_{i=1}^k \chi_{\mathbb{P}}(n + h_i) - \rho\Bigr) w_n,\] with \(\chi_{\mathbb{P}}\) the characteristic function of the primes. If \(S(N, \rho) > 0\) for all large \(N\), then infinitely many \(n\) have at least \(\lfloor \rho \rfloor + 1\) of the \(n+h_i\) prime: some summand is positive, and since \(w_n \ge 0\) that \(n\) satisfies \(\sum_i \chi_{\mathbb{P}}(n+h_i) > \rho\). Writing \[S_1(\lambda) = \sum_{N<n\leq 2N} w_n,\qquad S_2(\lambda) = \sum_{N<n\leq 2N} \Big(\sum_i \chi_{\mathbb{P}}(n+h_i)\Big) w_n,\] we have \[S(N, \rho) = S_2(\lambda) - \rho S_1(\lambda).\] One chooses \(w_n\) to make \(S_2(\lambda)/S_1(\lambda)\) large.
The sieve weights and the \(W\)-trick
Maynard’s weights are \[w_n \;=\; \Bigl(\sum_{d_i \mid n + h_i\ \forall i} \lambda(d_1, \dotsc, d_k)\Bigr)^{\!2}, \qquad \lambda : \mathbb{N}^k \to \mathbb{R},\] supported on \(\prod_i d_i \le R\), where \(R\) is a truncation parameter below the level of distribution \(\vartheta\). The one-dimensional choice \(\lambda(d_1, \dotsc, d_k) = \mu(d)\,F\bigl(\log(R/d)\bigr)\) with \(d = d_1 \dotsm d_k\) recovers the Selberg weights of Goldston–Pintz–Yıldırım; for these the level \(\vartheta < 1/2\) does not suffice to make \(S_2(\lambda) - \rho S_1(\lambda) > 0\) for any \(\rho \ge 1\), and it is the extra freedom of the \(k\)-dimensional \(\lambda\) that removes this obstruction. The \(\lambda(d_1, \dotsc, d_k)\) are built from a fixed profile \(F\) on a subset of \(\mathbb{R}^k\); the subset distinguishes the two bounds.
To handle small primes, we restrict \(w_n\) to be nonzero only for \(n\) in a single residue class \(v_0 \pmod W\), where \(W = \prod_{p \le D_0} p\) and \(D_0 = \log\log\log N\). By admissibility of \(\mathcal{H}\) and the Chinese remainder theorem there is a \(v_0\) with \(v_0 + h_i\) coprime to \(W\) for every \(i\), so no prime \(p \le D_0\) divides any \(n + h_i\).
The variational problem and the two cases
Write \(\mathcal{R}_k(s) := \{t \in \mathbb{R}_{\ge 0}^k : t_1 + \dotsb + t_k \le s\}\) for the standard simplex scaled to total mass \(s\), and \(\mathcal{R}_k := \mathcal{R}_k(1)\). Fix \(\varepsilon \ge 0\) and take \(F\) supported on \(\mathcal{R}_k(1+\varepsilon)\). With \(W\) the primorial of the \(W\)-trick, the sieve sums satisfy \[\begin{aligned} S_1(\lambda) &\;=\; \bigl(1 + o(1)\bigr)\, \frac{\varphi(W)^k}{W^{k+1}}\, N\, (\log R)^k\, I_k(F),\\ S_2(\lambda) &\;=\; \bigl(1 + o(1)\bigr)\, \frac{\log R}{\log N}\, \frac{\varphi(W)^k}{W^{k+1}}\, N\, (\log R)^{k} \sum_{m=1}^k J_{k,\varepsilon}^{(m)}(F), \end{aligned}\] where \(I_k(F) = \int_{\mathcal{R}_k(1+\varepsilon)} F^2\) and \(J_{k,\varepsilon}^{(m)}(F)\) is the \(L^2\) mass of the \(m\)th marginal \(\int F\,dt_m\) over the slice \(\mathcal{R}_{k-1}(1-\varepsilon)\) (Section 5 gives the precise definitions). Since \(\log R = (\vartheta/2 - \delta)\log N\), \[\frac{S_2(\lambda)}{S_1(\lambda)} \;\sim\; \Bigl(\frac{\vartheta}{2} - \delta\Bigr) M_{k,\varepsilon}, \qquad M_{k,\varepsilon} := \sup_{F} \frac{\sum_{m=1}^k J_{k,\varepsilon}^{(m)}(F)}{I_k(F)}.\] Provided \(\vartheta\) and \(\varepsilon\) satisfy a compatibility condition, \(S(N,\rho)\) is eventually positive when \(\rho < (\vartheta/2)\,M_{k,\varepsilon}\); two of the \(n+h_i\) are then prime for infinitely many \(n\), giving a gap at most the diameter of \(\mathcal{H}\), once \(M_{k,\varepsilon} > 2/\vartheta\).
The two bounds are the cases \(\varepsilon = 0\) and \(\varepsilon = 1/25\).
\(\varepsilon = 0\), \(k = 105\). Here \(\mathcal{R}_k(1+\varepsilon) = \mathcal{R}_k\) and \(M_{105,0} > 4\). Since Bombieri–Vinogradov gives every \(\vartheta < 1/2\) and \(4 > 2/\vartheta\) for \(\vartheta\) near \(1/2\), the criterion holds; with an admissible \(105\)-tuple of diameter \(600\) this gives \(H_1 \le 600\).
\(\varepsilon = 1/25\), \(k = 50\). Enlarging the support to \(\mathcal{R}_k(1+\varepsilon)\) raises the supremum to \(M_{50,1/25} > 4.0043\), exhibited by an explicit \(F\) from a finite basis of symmetric polynomials, for which the ratio is a quotient of rational quadratic forms verified by the certificate of Section 13. With an admissible \(50\)-tuple of diameter \(246\) this gives \(H_1 \le 246\).
The first case is unconditional given Bombieri–Vinogradov. The second also requires the finite numerical inequality \(M_{50,1/25} > 4.0043\) (Section 13), which we prove by exact computation.
Definitions and preliminaries
This section collects the objects and standard facts used throughout: the classical arithmetic functions, admissible sets, the level of distribution and the two external inputs (Bombieri–Vinogradov and the prime number theorem), the \(W\)-trick, and the truncation parameter \(R\) with the support condition it imposes on the sieve weights.
Arithmetic functions and standard notation
Notation 1 (Standard symbols). We write \(\mathbb{N} = \{1, 2, 3, \dotsc\}\) for the positive integers and \(\mathbb{P}= \{2, 3, 5, \dotsc\}\) for the primes. For integers \(a, b\) we write \((a, b) = \gcd(a, b)\) and \([a, b] = \operatorname{lcm}(a, b)\), and we call \(a\) and \(b\) coprime when \((a,b) = 1\). For real \(x\) we write \(\lfloor x \rfloor\) and \(\lceil x \rceil\) for the floor and the ceiling. Finally, \(\chi_{\mathbb{P}}: \mathbb{N} \to \{0, 1\}\) denotes the indicator function of the primes, so that \(\chi_{\mathbb{P}}(n) = 1\) if \(n \in \mathbb{P}\) and \(\chi_{\mathbb{P}}(n) = 0\) otherwise.
Notation 1 (Asymptotic notation). For real-valued functions \(f, g\), the statement \(f \ll g\) means \(|f| \le C|g|\) for some constant \(C\), which may depend on the tuple length \(k\) and on \(\mathcal{H}\) (both fixed throughout) but never on \(N\). A subscript, as in \(f \ll_A g\), records an additional permitted dependence of \(C\). We write \(f = O(g)\) for \(f \ll g\), and \(f = o(g)\) when \(f/g \to 0\) as \(N \to \infty\).
Definition 1 (Möbius function \(\mu\)). The Möbius function \(\mu : \mathbb{N} \to \{-1, 0, 1\}\) is defined by \[\mu(n) \;:=\; \begin{cases} 1 & \text{if } n = 1,\\ (-1)^r & \text{if } n = p_1 p_2 \dotsm p_r \text{ is a product of } r \text{ distinct primes},\\ 0 & \text{if } p^2 \mid n \text{ for some prime } p.\end{cases}\]
Definition 1 (Euler totient \(\varphi\)). The Euler totient function \(\varphi: \mathbb{N} \to \mathbb{N}\) is defined by \[\varphi(n) \;:=\; \#\{a \in \{1, 2, \dotsc, n\} : (a, n) = 1\}.\]
Definition 1 (\(r\)-fold divisor function \(\tau_r\)). For an integer \(r \ge 0\), the \(r\)-fold divisor function \(\tau_r : \mathbb{N} \to \mathbb{N}\) counts ordered factorisations into \(r\) positive-integer factors: \[\tau_r(n) \;:=\; \#\{(d_1, d_2, \dotsc, d_r) \in \mathbb{N}^r : d_1 d_2 \dotsm d_r = n\}.\] We abbreviate \(\tau := \tau_2\), the ordinary divisor-counting function \(\tau(n) = \#\{d \ge 1 : d \mid n\}\).
Definition 1 (The function \(g\)). Let \(g : \mathbb{N} \to \mathbb{N}\) be the multiplicative function determined on prime powers by \[g(p) \;:=\; p - 2, \qquad g(p^{a}) \;:=\; p^{a-2}(p-1)^2 \quad (a \ge 2),\] for every prime \(p\), together with \(g(1) = 1\).
Lemma 1 (Möbius divisor-sum identity). For every positive integer \(s\), \[\sum_{t \mid s} \mu(t) \;=\; \mathbf{1}[s = 1].\]
Uses: def_mobius
Proof. If \(s = 1\), the only divisor is \(t = 1\) and \(\mu(1) = 1\) by Definition 1, so the sum equals \(1\). Suppose \(s > 1\) and let \(p_1, \dotsc, p_r\) (\(r \ge 1\)) be the distinct primes dividing \(s\). By Definition 1, \(\mu(t) = 0\) unless \(t\) is squarefree, so only the \(2^r\) divisors of \(p_1 \dotsm p_r\) contribute, and such a divisor \(t = \prod_{i \in I} p_i\) contributes \((-1)^{|I|}\). Grouping by \(|I|\) gives \[\sum_{t \mid s} \mu(t) \;=\; \sum_{j = 0}^{r} \binom{r}{j} (-1)^{j} \;=\; (1 - 1)^{r} \;=\; 0\] by the binomial theorem, since \(r \ge 1\). ◻
Lemma 1 (Submultiplicativity of \(\tau_r\)). For every integer \(r \ge 0\) and all positive integers \(m, n\), \[\tau_r(mn) \;\le\; \tau_r(m)\, \tau_r(n),\] with equality when \((m, n) = 1\).
Uses: def_tau_r
Proof. By Definition 1, \(\tau_r\) is the \(r\)-fold Dirichlet convolution of the constant function \(\mathbf{1}\) with itself. Convolution preserves multiplicativity, whence \(\tau_r(mn) = \tau_r(m)\tau_r(n)\) whenever \((m,n) = 1\); this gives the asserted equality.
For the inequality in general, it suffices to produce an injection from the ordered \(r\)-fold factorisations of \(mn\) into the pairs consisting of an ordered \(r\)-fold factorisation of \(m\) and one of \(n\). Fix a factorisation \(mn = d_1 \dotsm d_r\). For each prime \(p\) distribute the exponent \(v_p(d_i)\) between \(m\) and \(n\) by the greedy rule: assign to the \(m\)-part the first \(\min(v_p(d_i), \text{remaining exponent of } p \text{ in } m)\) of the copies of \(p\) occurring in \(d_i\), processing \(i = 1, \dotsc, r\) in order, and assign the rest to the \(n\)-part. This produces factorisations \(m = a_1 \dotsm a_r\) and \(n = b_1 \dotsm b_r\) with \(d_i = a_i b_i\), and the original factorisation is recovered from the pair, so the assignment is injective. Hence \(\tau_r(mn) \le \tau_r(m)\tau_r(n)\). ◻
Lemma 1 (Equality \(\tau_r(n)^2 = \tau_{r^2}(n)\) on squarefree integers). For every integer \(r \ge 0\) and every squarefree positive integer \(n\), \[\tau_r(n)^2 \;=\; \tau_{r^2}(n).\]
Uses: def_tau_r
Proof. Both sides are multiplicative in \(n\) by Lemma 1, so for squarefree \(n = p_1 \dotsm p_s\) with distinct primes \(p_i\) it suffices to verify the identity at a single prime argument. An ordered \(r\)-fold factorisation of a prime \(p\) has exactly one factor equal to \(p\) and all others equal to \(1\), so \(\tau_r(p) = r\); likewise \(\tau_{r^2}(p) = r^2\). Hence \(\tau_r(p)^2 = r^2 = \tau_{r^2}(p)\), and multiplying over \(p_1, \dotsc, p_s\) gives \(\tau_r(n)^2 = \tau_{r^2}(n)\). (The hypothesis that \(n\) be squarefree is essential: for \(n = p^2\) one has \(\tau_r(p^2)^2 = \binom{r+1}{2}^2 < \binom{r^2+1}{2} = \tau_{r^2}(p^2)\) in general.) ◻
Proof uses: lem_tau_submultiplicative
Lemma 1 (Dirichlet convolution identity \(\varphi= g \ast \mathbf{1}\)). For every positive integer \(n\), \[\varphi(n) \;=\; \sum_{u \mid n} g(u).\]
Uses: def_g_func, def_totient
Proof. Both sides are multiplicative functions of \(n\): this is standard for \(\varphi\), and the divisor sum of the multiplicative function \(g\) (Definition 1) is again multiplicative. It therefore suffices to compare the two sides at a prime power \(n = p^{a}\) with \(a \ge 1\).
For \(a = 1\), \[\sum_{u \mid p} g(u) \;=\; g(1) + g(p) \;=\; 1 + (p - 2) \;=\; p - 1 \;=\; \varphi(p).\] For \(a \ge 2\), using \(g(p^{j}) = p^{j-2}(p-1)^2\) for \(2 \le j \le a\), \[\sum_{u \mid p^{a}} g(u) \;=\; 1 + (p-2) + (p-1)^2 \sum_{j=2}^{a} p^{j-2} \;=\; (p-1) + (p-1)^2 \cdot \frac{p^{a-1} - 1}{p - 1} \;=\; (p-1)\,p^{a-1},\] which is \(\varphi(p^{a})\). Multiplicativity now gives the identity for every \(n\). ◻
Lemma 1 (Euler product identity for \(\sum_{d \mid n} \mu(d)/d\)). For every positive integer \(n\), \[\sum_{d \mid n} \frac{\mu(d)}{d} \;=\; \frac{\varphi(n)}{n} \;=\; \prod_{p \mid n} \Bigl(1 - \frac{1}{p}\Bigr).\]
Uses: def_mobius, def_totient
Proof. The function \(n \mapsto \sum_{d \mid n} \mu(d)/d\) is the Dirichlet convolution of the multiplicative function \(d \mapsto \mu(d)/d\) with the constant function \(\mathbf{1}\), hence multiplicative. Evaluating at a prime power \(p^{a}\) with \(a \ge 1\) and using \(\mu(p^{j}) = 0\) for \(j \ge 2\) (Definition 1), \[\sum_{d \mid p^{a}} \frac{\mu(d)}{d} \;=\; \frac{\mu(1)}{1} + \frac{\mu(p)}{p} \;=\; 1 - \frac{1}{p}.\] Hence, writing \(n = \prod_{p \mid n} p^{a_p}\), the divisor sum equals \(\prod_{p \mid n}(1 - 1/p)\). On the other hand \(\varphi\) is multiplicative with \(\varphi(p^{a}) = p^{a}(1 - 1/p)\) (Definition 1), so \(\varphi(n)/n = \prod_{p \mid n}(1 - 1/p)\) as well. The two displayed equalities follow. ◻
Lemma 1 (Dirichlet expansion of \(n/\varphi(n)\)). For every positive integer \(n\), \[\frac{n}{\varphi(n)} \;=\; \sum_{e \mid n} \frac{\mu(e)^2}{\varphi(e)}.\]
Uses: def_mobius, def_totient
Proof. Both sides are multiplicative in \(n\); for the right-hand side this is because \(\mu^2\) and \(1/\varphi\) are multiplicative and the divisor sum of a multiplicative function is multiplicative. It therefore suffices to compare at a prime power \(n = p^{a}\) with \(a \ge 1\). Since \(\mu(p^{j})^2 = 0\) for \(j \ge 2\), \[\sum_{e \mid p^{a}} \frac{\mu(e)^2}{\varphi(e)} \;=\; \frac{1}{\varphi(1)} + \frac{1}{\varphi(p)} \;=\; 1 + \frac{1}{p-1} \;=\; \frac{p}{p-1} \;=\; \frac{p^{a}}{\varphi(p^{a})},\] the last equality because \(\varphi(p^{a}) = p^{a-1}(p-1)\). Multiplicativity gives the identity in general. ◻
Lemma 1 (Lower bound for \(\varphi\) in terms of \(\tau\)). For every positive integer \(n\), \[n \;\le\; \tau(n)\, \varphi(n), \qquad\text{equivalently}\qquad \frac{n}{\tau(n)} \;\le\; \varphi(n).\]
Uses: def_totient, def_tau_r
Proof. Write \(n = \prod_{i=1}^{r} p_i^{a_i}\) with distinct primes \(p_i\), so \(r = \omega(n)\) and \(\tau(n) = \prod_i (a_i + 1) \ge 2^{r}\). For every prime \(p \ge 2\) we have \(1 - 1/p \ge 1/2\), hence by Lemma 1 \[\frac{\varphi(n)}{n} \;=\; \prod_{i=1}^{r}\Bigl(1 - \frac{1}{p_i}\Bigr) \;\ge\; 2^{-r} \;\ge\; \frac{1}{\tau(n)}.\] Multiplying through by \(n\,\tau(n)\) gives \(n \le \tau(n)\varphi(n)\). ◻
Proof uses: lem_euler_product_mu_over_d
Lemma 1 (Union bound for finite sums). Let \(X\) be a finite index set, let \(\{A_x\}_{x \in X}\) be finite subsets of an ambient set, and write \(\mathbf{1}_{A}\) for the indicator function of \(A\). Then, for every point \(y\), \[\mathbf{1}_{\bigcup_{x \in X} A_x}(y) \;\le\; \sum_{x \in X} \mathbf{1}_{A_x}(y).\]
Proof. If \(y \notin \bigcup_x A_x\) both sides are \(0\). If \(y \in \bigcup_x A_x\), then \(y \in A_{x_0}\) for at least one \(x_0 \in X\), so the left-hand side equals \(1\) while the right-hand side is a sum of non-negative terms one of which equals \(1\); the inequality follows. ◻
Throughout the blueprint, Lemma 1 is applied in the following form. Given a sum \(\Sigma = \sum_{t \in \mathcal{T}} c(t)\) of non-negative terms and a covering \(\mathcal{T} = \bigcup_{x \in X} \mathcal{T}_x\) by finitely many sub-events, multiplying the displayed indicator inequality by \(c(t) \ge 0\) and summing yields \[\Sigma \;\le\; \sum_{x \in X}\, \sum_{t \in \mathcal{T}_x} c(t).\] The phrase “apply the union bound over the choices of \(x\)” denotes this passage.
Admissible sets
Definition 1 (Admissible set). Let \(k \in \mathbb{N}\). A finite set \(\mathcal{H} = \{h_1, \dotsc, h_k\} \subset \mathbb{Z}_{\ge 0}\) of \(k\) distinct non-negative integers is admissible if, for every prime \(p\), there exists an integer \(a_p\) with \[a_p \not\equiv h_i \pmod{p} \qquad \text{for every } i \in \{1, \dotsc, k\}.\]
Lemma 1 (Cardinality characterisation of admissibility). A finite set \(\mathcal{H} \subset \mathbb{Z}_{\ge 0}\) is admissible (Definition 1) if and only if \[\#\{h \bmod p : h \in \mathcal{H}\} \;<\; p \qquad \text{for every prime } p.\]
Uses: def_admissible
Proof. Fix a prime \(p\) and set \(S_p := \{h \bmod p : h \in \mathcal{H}\} \subseteq \mathbb{Z}/p\mathbb{Z}\). There exists \(a_p\) with \(a_p \not\equiv h_i \pmod p\) for every \(i\) if and only if \(S_p \ne \mathbb{Z}/p\mathbb{Z}\): given such an \(a_p\), its residue class is missing from \(S_p\); conversely, if \(S_p \subsetneq \mathbb{Z}/p\mathbb{Z}\), any representative of a class not in \(S_p\) will do. Since \(\mathbb{Z}/p\mathbb{Z}\) is finite of cardinality \(p\), the condition \(S_p \ne \mathbb{Z}/p\mathbb{Z}\) is equivalent to \(\#S_p < p\). Applying this for every prime \(p\) gives the claim. ◻
Lemma 1 (Admissibility is a finite check). A finite set \(\mathcal{H} \subset \mathbb{Z}_{\ge 0}\) is admissible (Definition 1) if and only if, for every prime \(p \le \#\mathcal{H}\), there exists \(a < p\) with \(h \not\equiv a \pmod p\) for every \(h \in \mathcal{H}\).
Uses: def_admissible
Proof. Admissibility trivially implies the stated condition, since the latter only restricts the primes \(p \le \#\mathcal{H}\) and, as in the proof of Lemma 1, one may always replace a missed residue by its representative in \(\{0, 1, \dotsc, p-1\}\).
Conversely, assume the condition and let \(p\) be an arbitrary prime. If \(p \le \#\mathcal{H}\) there is nothing to prove. If \(p > \#\mathcal{H}\), then the image of \(\mathcal{H}\) in \(\mathbb{Z}/p\mathbb{Z}\) has at most \(\#\mathcal{H} < p\) elements, so it is a proper subset of \(\mathbb{Z}/p\mathbb{Z}\) and some residue class is missed; by Lemma 1 this is exactly what admissibility requires at \(p\). Hence \(\mathcal{H}\) is admissible. ◻
Proof uses: lem_admissible_cardinality_char
Lemma 1 reduces admissibility — a condition on all primes — to a finite verification, so the admissibility of an explicit tuple is decidable.
Definition 1 (Diameter of \(\mathcal{H}\)). Let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) be a finite non-empty set of non-negative integers. The diameter of \(\mathcal{H}\) is \[\operatorname{diam}(\mathcal{H}) \;:=\; \max_{1 \le i, j \le k} (h_i - h_j).\]
Lemma 1 (Diameter as largest minus smallest). Let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) be a finite non-empty set of non-negative integers, labelled so that \(h_1 < h_2 < \dotsb < h_k\). Then \[\operatorname{diam}(\mathcal{H}) \;=\; h_k - h_1.\]
Uses: def_diam
Proof. The pair \((i, j) = (k, 1)\) is admissible in the maximum of Definition 1 and contributes \(h_k - h_1\), so \(\operatorname{diam}(\mathcal{H}) \ge h_k - h_1\). Conversely, for any \(i, j\) we have \(h_i \le h_k\) and \(h_j \ge h_1\), whence \(h_i - h_j \le h_k - h_1\); taking the maximum over \(i, j\) gives \(\operatorname{diam}(\mathcal{H}) \le h_k - h_1\). ◻
Prime counting and the level of distribution
Notation 1 (Prime-counting functions). For real \(x\), a modulus \(q \in \mathbb{N}\), a residue \(a\) modulo \(q\), and a subset \(S \subseteq \mathbb{N}\), set \[\pi_S(x) \;:=\; \#\{p \le x : p \in \mathbb{P},\ p \in S\}, \qquad \pi_S(x; q, a) \;:=\; \#\{p \le x : p \in \mathbb{P},\ p \in S,\ p \equiv a \!\!\pmod q\}.\] When \(S = \mathbb{N}\) we drop it from the notation and write \(\pi(x)\) and \(\pi(x; q, a)\); these are the classical prime-counting function and its analogue in an arithmetic progression. A real argument is always interpreted through its floor, so \(\pi_S(x) = \pi_S(\lfloor x \rfloor)\).
Definition 1 (Level of distribution). Let \(S \subseteq \mathbb{N}\), let \(\vartheta \in \mathbb{R}\), and let \(\Xi \in \mathbb{N}\) be an auxiliary modulus. We say that \(S\) has level of distribution \(\vartheta\) relative to \(\Xi\) if for every \(A \ge 1\) there is a constant \(c = c(A) > 0\) such that, for every real \(x \ge 3\), \[\sum_{\substack{1 \le q \le x^{\vartheta}\\ (q, \Xi) = 1}} \; \max_{\substack{a \bmod q\\ (a, q) = 1}} \; \Bigl| \pi_S(x; q, a) \;-\; \frac{\pi_S(x)}{\varphi(q)} \Bigr| \;\le\; \frac{c\,x}{(\log x)^{A}}.\] The unadorned phrase “the primes have level of distribution \(\vartheta\)” means that \(\mathbb{P}\) has level of distribution \(\vartheta\) relative to \(\Xi = 1\).
Uses: def_totient, not_pi
By Definition 1, the level of distribution depends on \(S\) only through its primes, so the hypothesis may be stated for a set \(S\) that is not itself a set of primes.
Lemma 1 (Insensitivity to intersecting with the primes). Let \(S \subseteq \mathbb{N}\), let \(\vartheta \in \mathbb{R}\) and let \(\Xi \in \mathbb{N}\). Then \(S \cap \mathbb{P}\) has level of distribution \(\vartheta\) relative to \(\Xi\) if and only if \(S\) does.
Proof. By Notation 1, both \(\pi_S\) and \(\pi_S(\cdot\,; q, a)\) count only elements of \(S\) that are prime, so \[\pi_{S \cap \mathbb{P}}(x) = \pi_S(x) \qquad\text{and}\qquad \pi_{S \cap \mathbb{P}}(x; q, a) = \pi_S(x; q, a)\] for all \(x, q, a\). The two instances of the inequality in Definition 1 are therefore literally the same, so each implies the other. ◻
Proof uses: not_pi
Theorem 1 (Bombieri–Vinogradov). For every \(\theta < 1/2\), the set \(\mathbb{N}\) has level of distribution \(\theta\) relative to \(\Xi = 1\).
Proof. This is the classical theorem of Bombieri and Vinogradov, taken as an external input; see Bombieri (1965), Vinogradov (1965), and the textbook treatment in Davenport’s Multiplicative Number Theory. No proof is given here. ◻
Lemma 1 (Bombieri–Vinogradov for the primes). Assume Bombieri–Vinogradov (Theorem 1). Then for every \(\theta < 1/2\) the set \(\mathbb{P}\) of primes has level of distribution \(\theta\) relative to \(\Xi = 1\).
Uses: thm_BV, def_level_of_distribution
Proof. Fix \(\theta < 1/2\). By Theorem 1 the set \(\mathbb{N}\) has level of distribution \(\theta\) relative to \(1\). Since \(\mathbb{P}= \mathbb{N} \cap \mathbb{P}\), Lemma 1 applied with \(S = \mathbb{N}\) transfers the property from \(\mathbb{N}\) to \(\mathbb{P}\). ◻
Proof uses: lem_level_prime_restriction
External analytic inputs
Two analytic results are quoted from the literature: Bombieri–Vinogradov (Theorem 1, above) and the prime number theorem, stated below. Bombieri–Vinogradov is a hypothesis of both main theorems and is never discharged. The prime number theorem is used once, to count the primes in the dyadic interval, and is quoted with a secondary term. The size bound for the primorial \(W\) uses Chebyshev-type bounds.
In a proof, “external input” means that the assertion is quoted from the literature and no proof is given here.
Theorem 2 (Prime number theorem with secondary term). There exist a constant \(C > 0\) and an integer \(N_0\) such that, for every integer \(N \ge N_0\), \[\Bigl| \pi(N) \;-\; \frac{N}{\log N} \Bigr| \;\le\; \frac{C\,N}{(\log N)^2}.\]
Uses: not_pi
Proof. This is the prime number theorem in the form \(\pi(x) = x/\log x + O\bigl(x/(\log x)^2\bigr)\), equivalently \(\psi(x) = x + O\bigl(x/\log x\bigr)\) after partial summation and the passage from \(\psi\) to \(\theta\) to \(\pi\). It is taken as an external input; no proof is given here. ◻
Prime counts in a dyadic interval
The sieve runs over the dyadic interval \(N < n \le 2N\). The two quantities below count the primes in this window and record their discrepancy in arithmetic progressions.
Notation 1 (Prime-counting in \((N, 2N]\)). \[X_N \;:=\; \sum_{N < n \le 2N} \chi_{\mathbb{P}}(n) \;=\; \#\{p \in \mathbb{P}: N < p \le 2N\}.\]
Uses: not_standard
Lemma 1 (\(X_N\) as a difference of prime-counting functions). For every \(N \ge 0\), \[X_N \;=\; \pi(2N) - \pi(N).\]
Proof. By Notation 1, \(\pi(x) = \sum_{n \le x} \chi_{\mathbb{P}}(n)\), the primes \(\le x\) being counted once each by the indicator \(\chi_{\mathbb{P}}\). The set of primes \(\le N\) is contained in the set of primes \(\le 2N\), so the two counts differ by exactly the primes in \((N, 2N]\): \[\pi(2N) - \pi(N) \;=\; \sum_{n \le 2N} \chi_{\mathbb{P}}(n) \;-\; \sum_{n \le N} \chi_{\mathbb{P}}(n) \;=\; \sum_{N < n \le 2N} \chi_{\mathbb{P}}(n) \;=\; X_N,\] the last equality being Notation 1. ◻
Theorem 3 (Prime number theorem for \((N, 2N]\)). There exist a constant \(C > 0\) and an integer \(N_0\) such that, for every integer \(N \ge N_0\), \[\Bigl| X_N \;-\; \frac{N}{\log N} \Bigr| \;\le\; \frac{C\,N}{(\log N)^2}.\]
Uses: not_XN
Proof. By Lemma 1, \(X_N = \pi(2N) - \pi(N)\). Let \(C_0\) and \(N_0'\) be as furnished by Theorem 2, and take \(N \ge \max(N_0', 2)\). Applying Theorem 2 at \(2N\) and at \(N\), and using \(\log(2N) \ge \log N\) to weaken the error term at \(2N\), \[\Bigl| \pi(2N) - \frac{2N}{\log(2N)} \Bigr| \le \frac{2C_0 N}{(\log N)^2}, \qquad \Bigl| \pi(N) - \frac{N}{\log N} \Bigr| \le \frac{C_0 N}{(\log N)^2}.\] It remains to compare the main terms. Since \(\log (2N) = \log 2 + \log N\), \[\frac{2N}{\log(2N)} - \frac{2N}{\log N} \;=\; -\,\frac{2N \log 2}{(\log 2 + \log N)\log N},\] whose absolute value is at most \(2N \log 2/(\log N)^2\) because \((\log 2 + \log N)\log N \ge (\log N)^2\). Combining, and writing \[X_N - \frac{N}{\log N} = \Bigl(\pi(2N) - \frac{2N}{\log 2N}\Bigr) - \Bigl(\pi(N) - \frac{N}{\log N}\Bigr) + \Bigl(\frac{2N}{\log 2N} - \frac{2N}{\log N}\Bigr),\] the triangle inequality gives the claim with \(C = 3C_0 + 2\log 2\). ◻
Proof uses: lem_XN_pi_diff, thm_PNT
Notation 1 (Local error \(E(N, q)\)). For \(q \in \mathbb{N}\), \[E(N, q) \;:=\; 1 \;+\; \sup_{(a, q) = 1} \Bigl| \sum_{\substack{N < n \le 2N\\ n \equiv a \!\pmod q}} \chi_{\mathbb{P}}(n) \;-\; \frac{1}{\varphi(q)} \sum_{N < n \le 2N} \chi_{\mathbb{P}}(n) \Bigr|,\] the supremum being over the residue classes \(a\) modulo \(q\) that are coprime to \(q\). The additive \(1\) is a convenience: it makes \(E(N,q) \ge 1\), so that \(E(N,q)\) may be used as a multiplicative majorant without a separate case for a vanishing discrepancy.
Uses: def_totient, not_XN
Lemma 1. Assume Bombieri–Vinogradov (Theorem 1). Let \(A > 0\) and \(\varepsilon \in (0, 1/2)\) be fixed. Then there is a constant \(C = C(A, \varepsilon) > 0\) such that, for every \(N \ge 1\), \[\sum_{\substack{1 \le q \le N\\ q < N^{1/2 - \varepsilon}}} \mu(q)^2 \, E(N, q) \;\le\; \frac{C\,N}{(\log N)^{A}}.\]
Uses: thm_BV, not_error, def_mobius
Proof. Write \(Q := \{q : 1 \le q \le N,\ q < N^{1/2-\varepsilon}\}\) for the range of moduli, and put \(\theta := 1/2 - \varepsilon\), so that \(0 < \theta < 1/2\) and \(Q \subseteq \{q : 1 \le q \le N^{\theta}\}\).
Step 1: bound \(E(N,q)\) by two cumulative discrepancies. Fix \(q\) and a residue \(a\) coprime to \(q\). The window count \(\sum_{N < n \le 2N,\, n \equiv a} \chi_{\mathbb{P}}(n)\) equals \(\pi(2N; q, a) - \pi(N; q, a)\), and \(\sum_{N < n \le 2N}\chi_{\mathbb{P}}(n) = X_N = \pi(2N) - \pi(N)\) by Lemma 1. Hence the quantity inside the supremum in Notation 1 equals \[\Bigl(\pi(2N; q, a) - \frac{\pi(2N)}{\varphi(q)}\Bigr) - \Bigl(\pi(N; q, a) - \frac{\pi(N)}{\varphi(q)}\Bigr),\] so by the triangle inequality \[E(N, q) \;\le\; 1 \;+\; D(2N, q) \;+\; D(N, q), \qquad D(x, q) := \max_{(a,q)=1}\Bigl|\pi(x; q, a) - \frac{\pi(x)}{\varphi(q)}\Bigr|.\]
Step 2: sum over \(q \in Q\). Since \(\mu(q)^2 \le 1\), summing Step 1 over \(q \in Q\) gives \[\sum_{q \in Q} \mu(q)^2 E(N,q) \;\le\; \# Q \;+\; \sum_{q \in Q} D(2N, q) \;+\; \sum_{q \in Q} D(N, q),\] and \(\#Q \le N^{1/2 - \varepsilon}\) by the definition of \(Q\).
Step 3: apply the level-of-distribution bound. Set \(A' := \max(A, 1)\). Since \(\theta < 1/2\), Lemma 1 says the primes have level of distribution \(\theta\) relative to \(\Xi = 1\), so Definition 1 supplies \(c = c(A') > 0\) with \(\sum_{q \le x^{\theta}} D(x, q) \le c\,x (\log x)^{-A'}\) for all \(x \ge 3\). All summands \(D(x,q)\) are non-negative, so restricting to \(q \in Q \subseteq \{q \le x^\theta\}\) (valid at \(x = N\) by the choice of \(\theta\), and at \(x = 2N\) since \(N^{\theta} \le (2N)^{\theta}\) for \(\theta \ge 0\)) yields, for \(N \ge 3\), \[\sum_{q \in Q} D(N, q) \le \frac{c\,N}{(\log N)^{A'}} \qquad\text{and}\qquad \sum_{q \in Q} D(2N, q) \le \frac{2c\,N}{(\log 2N)^{A'}} \le \frac{2c\,N}{(\log N)^{A'}}.\] Moreover \(\log N \ge 1\) for \(N \ge 3\), so \((\log N)^{-A'} \le (\log N)^{-A}\).
Step 4: absorb the trivial term and the small \(N\). The term \(\#Q \le N^{1/2 - \varepsilon}\) satisfies \(N^{1/2-\varepsilon} \le K N (\log N)^{-A}\) for a constant \(K = K(A, \varepsilon)\) and all \(N \ge 2\): indeed \((\log N)^{A} \ll_{A,\varepsilon} N^{1/2 + \varepsilon}\), and \(N^{1/2 - \varepsilon} \cdot N^{1/2+\varepsilon} = N\). Combining Steps 2–4 gives the asserted bound with a suitable constant for all \(N\) beyond an explicit threshold; enlarging the constant to accommodate the finitely many remaining \(N \ge 1\) (for which the left-hand side is a finite sum, and is empty when \(N = 1\)) gives the lemma. ◻
Proof uses: lem_XN_pi_diff, lem_BV_primes, def_level_of_distribution, not_XN
The \(W\)-trick
The primes are not equidistributed among the residue classes of a small modulus, and a sieve over \(N < n \le 2N\) inherits this defect in every local factor. The \(W\)-trick removes the obstruction by restricting \(n\) to a single residue class \(v_0\) modulo the product \(W\) of all small primes, chosen so that each shifted value \(n + h_i\) is coprime to \(W\). The cutoff defining “small” is a triple logarithm, large enough for the local factors to converge and small enough that \(W\) is negligible against any power of \(N\).
Definition 1 (The primorial cutoff \(D_0\)). For real \(N\), set \[D_0 \;:=\; \log\log\log N.\]
Definition 1 (The primorial \(W\)). Let \(D_0\) be as in Definition 1. Set \[W \;:=\; \prod_{\substack{p \le \lfloor D_0 \rfloor \\ p \in \mathbb{P}}} p,\] the product of all primes not exceeding \(\lfloor D_0 \rfloor\) (an empty product, equal to \(1\), when \(D_0 < 2\)).
Uses: def_D_0
Lemma 1 (Prime divisors of \(W\)). For every prime \(p\), \[p \mid W \iff p \le D_0.\]
Uses: def_W_trick, def_D_0
Proof. By Definition 1, \(W\) is a product of distinct primes, so a prime \(p\) divides \(W\) if and only if \(p\) is one of the factors, that is, if and only if \(p \le \lfloor D_0 \rfloor\). Since \(p\) is an integer, \(p \le \lfloor D_0 \rfloor\) is equivalent to \(p \le D_0\). ◻
Lemma 1 (\(W\) is squarefree). \(W\) is squarefree.
Uses: def_W_trick
Proof. By Definition 1, \(W\) is a product of distinct primes, each occurring to the first power. Hence no prime square divides \(W\), i.e. \(W\) is squarefree. ◻
Lemma 1 (Size of \(W\)). For all sufficiently large \(N\), \[W \;\le\; (\log\log N)^2.\]
Uses: def_W_trick, def_D_0
Proof. We use the elementary Chebyshev-type bound \(\prod_{p \le n} p \le 4^{n}\), valid for every integer \(n \ge 0\); it is proved by induction on \(n\) using the divisibility of \(\binom{2m+1}{m}\) by all primes in \((m+1, 2m+1]\) together with \(\binom{2m+1}{m} \le 4^{m}\). No appeal to the prime number theorem is required.
By Definition 1 and this bound, \[W \;=\; \prod_{p \le \lfloor D_0\rfloor} p \;\le\; 4^{\lfloor D_0 \rfloor} \;\le\; 4^{D_0}\] whenever \(D_0 \ge 0\). Now \(D_0 = \log\log\log N\) by Definition 1, so \(e^{D_0} = \log\log N\) and hence \[4^{D_0} \;=\; e^{D_0 \log 4} \;=\; \bigl(e^{D_0}\bigr)^{\log 4} \;=\; (\log\log N)^{\log 4}.\] Finally \(\log 4 \le 2\), so once \(N\) is large enough that \(\log\log N \ge 1\) — equivalently \(N \ge e^{e}\), which also forces \(D_0 \ge 0\) — we may raise the exponent from \(\log 4\) to \(2\) and conclude \(W \le (\log\log N)^{2}\). ◻
Definition 1 (Compatible residue class \(v_0\)). Let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) be a finite set of non-negative integers and let \(W\) be as in Definition 1. An integer \(v_0\) is compatible with \(\mathcal{H}\) modulo \(W\) if \[(v_0 + h_i,\, W) \;=\; 1 \qquad \text{for every } i \in \{1, \dotsc, k\}.\]
Uses: def_W_trick
Lemma 1 (Existence of \(v_0\)). Let \(\mathcal{H}\) be an admissible set (Definition 1) and let \(N\) be arbitrary. Then there exists \(v_0\) with \(0 \le v_0 < W\) that is compatible with \(\mathcal{H}\) modulo \(W\) (Definition 1).
Proof. By Definition 1, \(W\) is a product of distinct primes \(p \le \lfloor D_0 \rfloor\). For each such prime \(p\), admissibility of \(\mathcal{H}\) (Definition 1) supplies a residue \(a_p\) with \(a_p \not\equiv h_i \pmod p\) for every \(i\); equivalently, setting \(b_p := -a_p\), we have \(b_p + h_i \not\equiv 0 \pmod p\) for every \(i\).
The primes dividing \(W\) are pairwise coprime, so by the Chinese remainder theorem there is an integer \(v\) with \(v \equiv b_p \pmod p\) for every prime \(p \mid W\). Put \(v_0 := v \bmod W\), so \(0 \le v_0 < W\) and \(v_0 \equiv b_p \pmod p\) for each such \(p\).
It remains to check compatibility. Fix \(i\). Since \(W\) is squarefree (Lemma 1), \((v_0 + h_i, W) = 1\) holds if and only if no prime divisor of \(W\) divides \(v_0 + h_i\); and by Lemma 1 the prime divisors of \(W\) are exactly the primes \(p \le D_0\), which are the primes handled above. For such a \(p\) we have \(v_0 + h_i \equiv b_p + h_i \not\equiv 0 \pmod p\) by the choice of \(b_p\). Hence \((v_0 + h_i, W) = 1\) for every \(i\), i.e. \(v_0\) is compatible with \(\mathcal{H}\) modulo \(W\). ◻
Proof uses: lem_W_prime_dvd, lem_W_squarefree
Notation 1 (Standing notation). We fix throughout: an integer \(k \ge 2\); an admissible set \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) with \(h_1 < \dotsb < h_k\); the parameters \(D_0\) and \(W\) of Definitions 1 and 1; and a residue \(v_0\) compatible with \(\mathcal{H}\) modulo \(W\) (Definition 1). All constants implicit in \(O(\cdot)\), \(\ll\) and \(o(\cdot)\) may depend on \(k\) and on \(\mathcal{H}\), but never on \(N\).
The truncation parameter \(R\) and the sieve support
Definition 1 (The auxiliary parameter \(\delta\)). Let \(\vartheta\) be a level of distribution (Definition 1). Throughout the blueprint, \(\delta\) denotes a fixed real number in the open interval \[\delta \;\in\; \bigl(0,\, \vartheta/2\bigr),\] which may be chosen arbitrarily small. The proofs treat \(\delta\) as a fixed positive constant, and the dependence of error terms on \(\delta\) is suppressed in the asymptotic notation of Notation 1.
Definition 1 (The truncation \(R\)). Let \(\delta\) be as in Definition 1. The sieve truncation parameter is \[R \;:=\; N^{\vartheta/2 - \delta}.\]
Uses: def_delta
Lemma 1 (\(R^2 W\) lies below \(N^{1/2}\)). Let \(\vartheta\) and \(\delta\) be as in Definitions 1 and 1, let \(R\) be as in Definition 1, and let \(\eta\) be a real number with \[\eta \;<\; \tfrac12 - (\vartheta - 2\delta).\] Then, for all sufficiently large \(N\), \[R^{2}\, W \;\le\; N^{1/2 - \eta}.\]
Proof. Set \(\gamma := \tfrac12 - \eta - (\vartheta - 2\delta)\), which is positive by hypothesis. By Definition 1, \[R^{2} \;=\; N^{\vartheta - 2\delta},\] so it suffices to show that \(W \le N^{\gamma}\) for all large \(N\), since then \(R^{2}W \le N^{\vartheta - 2\delta + \gamma} = N^{1/2 - \eta}\).
By Lemma 1, \(W \le (\log\log N)^{2}\) for all large \(N\), and \(\log\log N \le \log N\) for \(N \ge e\). Hence \(W \le (\log N)^{2}\) for all large \(N\). Since \((\log N)^{2} = o(N^{\gamma})\) for every \(\gamma > 0\) — powers of the logarithm are dominated by every positive power of \(N\) — we obtain \(W \le N^{\gamma}\) for all large \(N\), as required. ◻
Proof uses: lem_W_size
The support condition below is imposed on every sieve weight. It makes the divisor sums finite, keeps the moduli below the Bombieri–Vinogradov range through Lemma 1, and decouples the sieve from the small primes absorbed into \(W\).
Definition 1 (Permissible support at a truncation level). Let \(T\) be a positive integer, the truncation level. A function \(\lambda : \mathbb{N}^{k} \to \mathbb{R}\) has permissible support at level \(T\) if \(\lambda(d_1, \dotsc, d_k) = 0\) whenever one of the following three conditions fails:
\(\prod_{i=1}^{k} d_i \le T\);
\(\bigl(\prod_{i=1}^{k} d_i,\, W\bigr) = 1\);
\(\prod_{i=1}^{k} d_i\) is squarefree.
Two levels occur below. The development on the standard simplex takes \(T = \lfloor R \rfloor\), and the development on the enlarged simplex takes \(T = \lfloor R^{1+\varepsilon} \rfloor\); the enlargement buys exactly the longer divisors that the larger level admits. Where the level is clear from context we say simply permissible support.
Uses: def_R, def_W_trick
Lemma 1 (Permissibility passes to divisors). Let \(\mathbf{r} = (r_1, \dotsc, r_k)\) satisfy the three conditions (i)–(iii) of Definition 1, and let \(\mathbf{d} = (d_1, \dotsc, d_k)\) satisfy \(d_i \mid r_i\) for every \(i\). Then \(\mathbf{d}\) satisfies (i)–(iii) as well.
Uses: def_adm_support
Proof. Since \(d_i \mid r_i\) for every \(i\), we have \(\prod_i d_i \mid \prod_i r_i\). Because \(\prod_i r_i\) is squarefree it is in particular non-zero, so \(\prod_i d_i \le \prod_i r_i \le R\), giving (i). A divisor of an integer coprime to \(W\) is itself coprime to \(W\), giving (ii), and a divisor of a squarefree integer is squarefree, giving (iii). ◻
Notation 1 (Singular series). For a density \(\gamma : \mathbb{N} \to \mathbb{R}\) the associated singular series is \[\mathfrak{S}(\gamma) \;:=\; \prod_{p \in \mathbb{P}} \frac{1 - 1/p}{1 - \gamma(p)/p},\] the product being taken over all primes.
The sieve datum, the sieve weights, and the auxiliary \(y\)-variables
The sieve datum
The sieve estimates below — the Mertens-type partial summations, the evaluation of the local densities, the singular products — depend on the arithmetic function \(\gamma_g\) only through a short list of properties: multiplicativity, a uniform gap between \(\gamma(p)\) and \(p\), closeness of \(\gamma(p)\) to \(1\) away from a fixed squarefree modulus, and a two-sided Mertens bound on the weighted prime log-sums. We collect this list as a single object, the sieve datum, consisting of the density \(\gamma\), the modulus \(V\), and the four constants \(A_1, A_2, A_3, L\). Every statement below is proved for an arbitrary sieve datum and applied twice: once to the \(W\)-tricked density \(\gamma_g\), and once, through the same lemmas with a different modulus, in the \(r_i\)-coordinate analysis of \(S_2^{(m)}\).
Definition 1 (Mertens deviation of a weight). Let \(f : \mathbb{N} \to \mathbb{R}\) and let \(2 \le w \le z\) be real. The Mertens deviation of \(f\) on \([w, z]\) is \[\Delta(f; w, z) \;:=\; \sum_{\substack{w \le p \le z \\ p \in \mathbb{P}}} \frac{f(p)\,\log p}{p} \;-\; \log(z/w).\]
Definition 1 (Density). A density is a function \(\gamma : \mathbb{N} \to \mathbb{R}\) such that
\(\gamma(n) \ge 0\) for every \(n\);
\(\gamma(1) = 1\);
\(\gamma(mn) = \gamma(m)\,\gamma(n)\) whenever \((m, n) = 1\);
\(\gamma(p) < p\) for every prime \(p\).
Definition 1 (Sieve datum). A sieve datum is a tuple \(\mathcal{S} = (\gamma, V, A_1, A_2, A_3, L)\) consisting of a function \(\gamma : \mathbb{N} \to \mathbb{R}\), an integer \(V \ge 1\), and real numbers \(A_1, A_2, A_3, L\), subject to the following axioms.
Density. \(\gamma\) is a density in the sense of Definition 1.
Growth. \(0 < A_1 < 1\) and \(\gamma(p)/p \le 1 - A_1\) for every prime \(p\).
Modulus. \(V\) is squarefree, and \(\gamma(p) = 0\) for every prime \(p \mid V\).
Unit density off the modulus. \(A_3 \ge 0\) and \(|\gamma(p) - 1| \le A_3/p\) for every prime \(p \nmid V\).
Mertens. \(A_2 > 0\), \(L \ge 0\), and \(-L \le \Delta(\gamma; w, z) \le A_2\) for all real \(2 \le w \le z\), with \(\Delta\) as in Definition 1.
We call \(\gamma\) the density, \(V\) the modulus, and \(A_1, A_2, A_3, L\) the constants of \(\mathcal{S}\).
Axiom (D1) is stated separately from the sieve datum because the same notion of density is also imposed level by level on the \(N\)-indexed family of \(W\)-tricked densities from which the constants \(A_2\) and \(L\) are extracted; the two uses share one definition rather than duplicating four axioms. Note also that (D2) forces \(\gamma(p) \le (1 - A_1)p\) uniformly in \(p\), which is strictly stronger than the pointwise inequality \(\gamma(p) < p\) of (D1)(iv); both are recorded because the weaker one is what the construction of \(g_*\) needs, and the stronger one is what the analytic estimates need.
Definition 1 (The derived multiplicative function \(g_*\)). Let \(\mathcal{S} = (\gamma, V, A_1, A_2, A_3, L)\) be a sieve datum (Definition 1). The derived function \(g_* : \mathbb{N} \to \mathbb{R}\) is the totally multiplicative function determined by its prime values \[g_*(p) \;:=\; \frac{\gamma(p)}{p - \gamma(p)};\] explicitly, \(g_*(n) := \prod_{p} g_*(p)^{v_p(n)}\), where \(v_p(n)\) is the exponent of \(p\) in \(n\), so that \(g_*(1) = 1\).
Uses: def_sieve_datum
Lemma 1 (\(g_*\) is well defined and non-negative). Let \(\mathcal{S}\) be a sieve datum and \(p\) a prime. Then \(p - \gamma(p) > 0\), and \(g_*(p) \ge 0\).
Uses: def_sieve_datum, def_g_star
Proof. Axiom (D1)(iv) of Definition 1 gives \(\gamma(p) < p\), i.e. \(p - \gamma(p) > 0\); hence the quotient defining \(g_*(p)\) makes sense. (Axiom (D2) sharpens this to \(p - \gamma(p) \ge A_1 p\), a bound uniform in \(p\), which is what the analytic applications use.) Since \(\gamma(p) \ge 0\) by (D1)(i) and the denominator is positive, \(g_*(p) = \gamma(p)/(p - \gamma(p)) \ge 0\). ◻
Because \(g_*\) is totally multiplicative and non-negative at every prime, \(g_*(n) \ge 0\) for every \(n\), and on a squarefree \(n\) one has \(g_*(n) = \prod_{p \mid n} g_*(p)\). We use both facts freely below.
Definition 1 (The convolution summand \(h\)). Let \(\mathcal{S}\) be a sieve datum with derived function \(g_*\) (Definition 1). Define \(h : \mathbb{N} \to \mathbb{R}\) by \[h(d) \;:=\; \mu(d)^2\, g_*(d).\]
Uses: def_g_star, def_mobius
Lemma 1 (\(h\) is supported on integers coprime to the modulus). Let \(\mathcal{S} = (\gamma, V, A_1, A_2, A_3, L)\) be a sieve datum. If \(d \ge 1\) satisfies \((d, V) > 1\), then \(h(d) = 0\).
Uses: def_sieve_datum, def_h_conv, def_g_star
Proof. Since \((d, V) > 1\), the integer \((d, V)\) has a prime divisor \(p\); then \(p \mid d\) and \(p \mid V\). By axiom (D3) of Definition 1, \(\gamma(p) = 0\); by Lemma 1 the denominator \(p - \gamma(p)\) is strictly positive, so Definition 1 gives \(g_*(p) = 0/(p - \gamma(p)) = 0\). Write \(d = p\,c\) with \(c \ge 1\). Total multiplicativity of \(g_*\) gives \(g_*(d) = g_*(p)\, g_*(c) = 0\), whence \(h(d) = \mu(d)^2\, g_*(d) = 0\). ◻
Proof uses: lem_g_star_well_defined
Thus \(h\) is a non-negative arithmetic function supported on the squarefree integers coprime to \(V\), and its partial sums are the input to the partial-summation lemma below. The next function records the deviation of \(g_*\) from the model density \(1/(p-1)\) at each prime not dividing the modulus.
Definition 1 (The convolution defect \(b\)). Let \(\mathcal{S} = (\gamma, V, A_1, A_2, A_3, L)\) be a sieve datum with derived function \(g_*\). Define \(b : \mathbb{N} \to \mathbb{R}\) by \[b(n) \;:=\; \begin{cases} \displaystyle\prod_{p \mid n} \Bigl(g_*(p) - \frac{1}{p-1}\Bigr), & \text{if $n$ is squarefree and $(n, V) = 1$},\\[1.2em] 0, & \text{otherwise}, \end{cases}\] the product being over the distinct primes dividing \(n\) (so \(b(1) = 1\)).
Uses: def_sieve_datum, def_g_star
By construction \(b\) is multiplicative, vanishes off the squarefree integers coprime to \(V\), and takes the value \[b(p) \;=\; g_*(p) - \frac{1}{p-1} \;=\; \frac{\gamma(p)}{p - \gamma(p)} - \frac{1}{p - 1} \;=\; \frac{p\,(\gamma(p) - 1)}{(p - \gamma(p))(p-1)}\] at a prime \(p \nmid V\); axiom (D4) therefore makes \(b(p)\) of size \(O(1/p^2)\).
We now record the density to which the construction is applied.
Definition 1 (The \(W\)-tricked density \(\gamma_g\)). Let \(V \ge 1\). Define \(\gamma_g^{(V)} : \mathbb{N} \to \mathbb{R}\) to be the totally multiplicative function with prime values \[\gamma_g^{(V)}(p) \;:=\; \begin{cases} 0, & p \mid V,\\[0.3em] \dfrac{p}{p-1}, & p \nmid V, \end{cases}\] that is, \(\gamma_g^{(V)}(n) := \prod_p \gamma_g^{(V)}(p)^{v_p(n)}\), with \(v_p(n)\) the exponent of \(p\) in \(n\).
Lemma 1 (The Maynard sieve datum). Let \(V \ge 1\) be squarefree with \(2 \mid V\), and suppose \(A_2 > 0\) and \(L \ge 0\) are such that \(-L \le \Delta\bigl(\gamma_g^{(V)}; w, z\bigr) \le A_2\) for all real \(2 \le w \le z\). Let \(\mathcal{S}_V\) be the six-tuple \[\mathcal{S}_V = \Bigl(\gamma_g^{(V)},\ V,\ A_1 = \tfrac{1}{2},\ A_2,\ A_3 = 2,\ L\Bigr).\] Then \(\mathcal{S}_V\) is a sieve datum in the sense of Definition 1.
Proof. Write \(\gamma := \gamma_g^{(V)}\). Axiom (D5) is the hypothesis, and the squarefreeness of \(V\) together with \(\gamma(p) = 0\) for \(p \mid V\) (immediate from Definition 1) is axiom (D3). We check the remaining axioms.
(D1) Density. Every prime value \(\gamma(p)\) is non-negative — it is either \(0\) or \(p/(p-1) > 0\) — so the product over the prime factorisation defining \(\gamma(n)\) is non-negative, and \(\gamma(1) = 1\) as an empty product. Total multiplicativity of \(\gamma\) over prime factorisations gives \(\gamma(mn) = \gamma(m)\gamma(n)\) for all coprime \(m, n\) (indeed for all nonzero \(m, n\)). For (iv), if \(p \mid V\) then \(\gamma(p) = 0 < p\). If \(p \nmid V\), then \(p \ne 2\) because \(2 \mid V\), so \(p \ge 3\); and \(p/(p-1) < p\) is equivalent to \(1 < p - 1\), i.e. to \(p > 2\), which holds.
(D2) Growth with \(A_1 = 1/2\). Certainly \(0 < 1/2 < 1\). If \(p \mid V\) then \(\gamma(p)/p = 0 \le 1/2\). If \(p \nmid V\) then \(p \ge 3\) as above, and \(\gamma(p)/p = 1/(p-1) \le 1/2\).
(D4) Unit density with \(A_3 = 2\). Clearly \(2 \ge 0\). Let \(p \nmid V\). Then \[|\gamma(p) - 1| \;=\; \Bigl|\frac{p}{p-1} - 1\Bigr| \;=\; \frac{1}{p-1},\] and \(1/(p-1) \le 2/p\) is equivalent to \(p \le 2(p-1)\), i.e. to \(p \ge 2\), which holds for every prime. ◻
In the application \(V\) is the primorial modulus \(W\) of the \(W\)-trick (Definition 1), which is squarefree and even as soon as \(D_0(N) \ge 2\); the constants \(A_2\) and \(L\) are then supplied, uniformly in \(N\), by the Mertens estimates for the family \(\gamma_g\).
Tensor sieve coefficients
Notation 1 (Boldface \(k\)-tuple convention). A boldface symbol such as \(\mathbf{d}\) denotes the \(k\)-tuple of positive integers \((d_1, \dotsc, d_k)\), and we use the following conventions throughout.
\(\lambda(\mathbf{d})\) denotes the value \(\lambda(d_1, \dotsc, d_k)\). The bare symbol \(\lambda\) denotes the function \(\mathbb{N}^k \to \mathbb{R}\) itself (as when it is passed as an argument to \(w_n\), \(S_1\), \(S_2\) or \(S_2^{(m)}\)); we never write a subscripted form such as \(\lambda_{\mathbf{d}}\).
\(\sum_{\mathbf{d}}\) denotes the iterated sum \(\sum_{d_1}\sum_{d_2}\dotsb\sum_{d_k}\), each variable ranging over \(\mathbb{N}\), and \(\prod_i d_i\) the product \(d_1 d_2 \dotsm d_k\).
\(\sum_{d_i \mid n + h_i\,\forall i}\) abbreviates \(\sum_{d_1 \mid n+h_1}\sum_{d_2 \mid n+h_2}\dotsb\sum_{d_k \mid n+h_k}\).
Analogous conventions apply to \(\mathbf{e}\), \(\mathbf{r}\), \(\mathbf{u}\), \(\mathbf{a}\), \(\mathbf{b}\).
All the sieve coefficient functions occurring below are finitely supported: the set \(\{\mathbf{d} : \lambda(\mathbf{d}) \ne 0\}\) is finite. Every \(\mathbf{d}\)-sum written without further qualification is therefore a finite sum, and no convergence question arises.
The sieve coefficients are indexed by \(k\)-tuples of divisors, one per shift \(h_i\), subject to three conditions: a truncation on the product of the divisors, coprimality to the \(W\)-trick modulus, and squarefreeness of the product. The truncation level is carried as a parameter rather than fixed to \(R\): Instantiation I runs at level \(R\) and Instantiation II at level \(R^{1+\varepsilon}\), and every lemma below is stated once, at an arbitrary level.
Two consequences of (iii) are used constantly and without comment: each individual \(d_i\) is squarefree (being a divisor of the squarefree \(\prod_i d_i\)), and the \(d_i\) are pairwise coprime. Both instantiations take \(\lambda\) with permissible support: Instantiation I at level \(T = \lfloor R\rfloor\), Instantiation II at level \(T = \lfloor R^{1+\varepsilon}\rfloor\).
The sieve weights and the sums \(S_1(\lambda)\), \(S_2(\lambda)\)
Definition 1 (Sieve weight). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported and let \(n \ge 1\) be an integer. The sieve weight at \(n\) is \[w_n \;:=\; \Bigl(\sum_{d_i \mid n + h_i\ \forall i} \lambda(d_1, \dotsc, d_k)\Bigr)^{\!2}.\]
Uses: not_bold_tuple
Lemma 1 (Non-negativity of \(w_n\)). For every finitely supported \(\lambda : \mathbb{N}^k \to \mathbb{R}\) and every integer \(n \ge 1\), we have \(w_n \ge 0\).
Uses: def_weight
Proof. By Definition 1, \(w_n\) is the square of the real number \(\sum_{d_i \mid n+h_i\ \forall i} \lambda(\mathbf{d})\), and squares of real numbers are non-negative. ◻
Non-negativity of \(w_n\) is the only property of the weight used in the passage from positivity of the sieve sums to primes; everything else about \(w_n\) enters through the asymptotic evaluation of \(S_1\) and \(S_2\).
Definition 1 (The sum \(S_1(\lambda)\)). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported, with sieve weight \(w_n\) as in Definition 1. Set \[S_1(\lambda) \;:=\; \sum_{\substack{N < n \le 2N \\ n \equiv v_0 \!\!\pmod W}} w_n .\]
Uses: def_weight, def_W_trick, def_v0_compatible
Definition 1 (The sum \(S_2(\lambda)\)). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported, with sieve weight \(w_n\) as in Definition 1. Set \[S_2(\lambda) \;:=\; \sum_{\substack{N < n \le 2N \\ n \equiv v_0 \!\!\pmod W}} \Bigl(\sum_{i=1}^{k} \chi_{\mathbb{P}}(n + h_i)\Bigr) w_n ,\] so that each \(n\) is weighted by the number of indices \(i\) for which \(n + h_i\) is prime.
Uses: def_weight, def_W_trick, def_v0_compatible
Definition 1 (The sum \(S_2^{(m)}(\lambda)\)). Let \(m \in \{1, \dotsc, k\}\) and let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported, with sieve weight \(w_n\) as in Definition 1. Set \[S_2^{(m)}(\lambda) \;:=\; \sum_{\substack{N < n \le 2N \\ n \equiv v_0 \!\!\pmod W}} \chi_{\mathbb{P}}(n + h_m)\, w_n .\]
Uses: def_weight, def_W_trick, def_v0_compatible
Lemma 1 (Decomposition of \(S_2(\lambda)\)). For every finitely supported \(\lambda : \mathbb{N}^k \to \mathbb{R}\), \[S_2(\lambda) \;=\; \sum_{m=1}^{k} S_2^{(m)}(\lambda).\]
Proof. Fix \(n\) in the index set of Definition 1. Distributing the finite inner sum across the product with \(w_n\) gives \[\Bigl(\sum_{i=1}^{k} \chi_{\mathbb{P}}(n + h_i)\Bigr) w_n \;=\; \sum_{i=1}^{k} \chi_{\mathbb{P}}(n + h_i)\, w_n .\] Substituting this into Definition 1 and interchanging the two finite summations — over \(n \in (N, 2N]\) with \(n \equiv v_0 \pmod W\), and over \(i \in \{1, \dotsc, k\}\) — yields \[S_2(\lambda) \;=\; \sum_{i=1}^{k}\ \sum_{\substack{N < n \le 2N \\ n \equiv v_0 \!\!\pmod W}} \chi_{\mathbb{P}}(n + h_i)\, w_n .\] By Definition 1 the inner sum is exactly \(S_2^{(i)}(\lambda)\); renaming the index \(i\) to \(m\) gives the claim. ◻
Definition 1 (The sieve difference \(S(\lambda, \rho)\)). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported and let \(\rho > 0\) be real. Set \[S(\lambda, \rho) \;:=\; S_2(\lambda) \;-\; \rho\, S_1(\lambda).\]
Unwinding Definitions 1 and 1, the sieve difference is the single sum \[S(\lambda, \rho) \;=\; \sum_{\substack{N < n \le 2N \\ n \equiv v_0 \!\!\pmod W}} \Bigl(\sum_{i=1}^{k} \chi_{\mathbb{P}}(n + h_i) \;-\; \rho\Bigr) w_n ,\] whose summands are the weights \(w_n \ge 0\) multiplied by the excess of the prime count over \(\rho\). Positivity of \(S(\lambda,\rho)\) therefore cannot come from the weights alone; it forces the excess to be positive somewhere. This is the GPY reduction, and it is the sole point at which the sieve produces primes.
Lemma 1. Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported and let \(\rho > 0\). If \(S(\lambda, \rho) > 0\), then there exists an integer \(n\) with \(N < n \le 2N\) and \(n \equiv v_0 \pmod W\) such that at least \(\lfloor \rho + 1 \rfloor\) of \(n + h_1, \dotsc, n + h_k\) are prime.
Uses: def_S_sum, def_S1, def_S2, def_weight, def_W_trick, lem_exists_v0
Proof. Write \(c(n) := \#\{i \in \{1, \dotsc, k\} : n + h_i \in \mathbb{P}\} = \sum_{i=1}^k \chi_{\mathbb{P}}(n+h_i)\), so that, as computed above, \[S(\lambda, \rho) \;=\; \sum_{\substack{N < n \le 2N \\ n \equiv v_0 \!\!\pmod W}} \bigl(c(n) - \rho\bigr) w_n .\] Suppose, for contradiction, that \(c(n) \le \rho\) for every \(n\) in the index set. By Lemma 1 each \(w_n \ge 0\), so each summand \((c(n) - \rho) w_n\) is non-positive, and therefore \(S(\lambda, \rho) \le 0\), contradicting the hypothesis. Hence there is an \(n\) in the index set with \(c(n) > \rho\).
It remains to convert \(c(n) > \rho\) into \(c(n) \ge \lfloor \rho + 1\rfloor\). Since \(c(n)\) is an integer and \(c(n) > \rho\), we have \(c(n) \ge \lfloor \rho \rfloor + 1\) when \(\rho \notin \mathbb{Z}\), and \(c(n) \ge \rho + 1\) when \(\rho \in \mathbb{Z}\); in either case \(c(n) \ge \lfloor \rho + 1 \rfloor\), because \(\lfloor \rho + 1\rfloor = \lfloor \rho\rfloor + 1\). ◻
Proof uses: lem_wn_nonneg
The auxiliary \(y\)-variables
The coefficients \(\lambda(\mathbf{d})\) are natural for the divisor-sum manipulations of \(S_1\) and \(S_2\), but ill-suited as variational unknowns: the divisibility relations couple them, and the diagonal form of the sieve sums is not visible in them. The change of variables \(\lambda \leftrightarrow y\) below is a linear, involutive Möbius-type substitution which diagonalises \(S_1\) and, after a further substitution \(y \rightsquigarrow y^{(m)}\), also \(S_2^{(m)}\). We give here the definitions and the inversion identities; the asymptotic evaluation of \(S_1\) and \(S_2^{(m)}\) in terms of \(y\) and \(y^{(m)}\) is carried out later.
Definition 1 (The \(y\)-transform \(y_\lambda\) of \(\lambda\)). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported. Define \(y_\lambda : \mathbb{N}^k \to \mathbb{R}\) by \[y_\lambda(r_1, \dotsc, r_k) \;:=\; \Bigl(\prod_{i=1}^{k} \mu(r_i)\, \varphi(r_i)\Bigr) \sum_{\substack{\mathbf{d} \in \mathbb{N}^k \\ r_i \mid d_i \text{ and } d_i \text{ squarefree } \forall i}} \frac{\lambda(d_1, \dotsc, d_k)}{\prod_{i=1}^{k} d_i}.\] The \(\mathbf{d}\)-sum is finite because \(\lambda\) is finitely supported, and \(y_\lambda\) is again finitely supported. We write \(y := y_\lambda\) and call its values the \(y\)-variables of \(\lambda\).
Uses: def_mobius, def_totient, not_bold_tuple
The squarefreeness restriction on the summation variables is vacuous in every application: if \(\lambda\) has permissible support at some level, then each \(d_i\) in its support is squarefree by Definition 1(iii), and the sum reduces to the unrestricted \(\sum_{r_i \mid d_i\,\forall i}\). Carrying the restriction in the definition lets the inversion identities below hold for arbitrary finitely supported \(\lambda\), with no support hypothesis; the map \(\lambda \mapsto y_\lambda\) is \(\mathbb{R}\)-linear.
Lemma 1 (The \(y\)-variables inherit permissible support). Let \(T \ge 1\) and let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) have permissible support at level \(T\) (Definition 1). Then \(y_\lambda\) has permissible support at level \(T\).
Uses: def_adm_support, def_y_from_lambda
Proof. Suppose \(y_\lambda(\mathbf{r}) \ne 0\). Then some summand of the defining sum in Definition 1 is nonzero, so there is a \(\mathbf{d}\) with \(\lambda(\mathbf{d}) \ne 0\) and \(r_i \mid d_i\) for every \(i\). It therefore suffices to show that the three conditions of Definition 1 are inherited from \(\mathbf{d}\) to any coordinatewise divisor \(\mathbf{r}\) of \(\mathbf{d}\). That is exactly Lemma 1, applied at level \(T\).
In detail: since \(r_i \mid d_i\) for each \(i\), we have \(\prod_i r_i \mid \prod_i d_i\). As all \(d_i\) are nonzero, \(\prod_i d_i \ge 1\) and hence \(\prod_i r_i \le \prod_i d_i \le T\), which is (i). A divisor of an integer coprime to \(W\) is coprime to \(W\), giving (ii); and a divisor of a squarefree integer is squarefree, giving (iii). ◻
Proof uses: lem_adm_support_dvd_closed
Definition 1 (The Möbius transform \(\lambda_y\) of \(y\)). Let \(y : \mathbb{N}^k \to \mathbb{R}\) be finitely supported. Define \(\lambda_y : \mathbb{N}^k \to \mathbb{R}\) by \[\lambda_y(d_1, \dotsc, d_k) \;:=\; \Bigl(\prod_{i=1}^{k} \mu(d_i)\, d_i\Bigr) \sum_{\substack{\mathbf{r} \in \mathbb{N}^k \\ d_i \mid r_i \text{ and } r_i \text{ squarefree } \forall i}} \frac{y(r_1, \dotsc, r_k)}{\prod_{i=1}^{k} \varphi(r_i)}.\] The \(\mathbf{r}\)-sum is finite because \(y\) is finitely supported, and \(\lambda_y\) is again finitely supported.
Uses: def_mobius, def_totient, not_bold_tuple
The transform \(y \mapsto \lambda_y\) is \(\mathbb{R}\)-linear, and the argument of Lemma 1 applies verbatim (with the roles of \(\mathbf{d}\) and \(\mathbf{r}\) exchanged) to show that it too preserves permissible support at any level \(T\). The next two lemmas identify the two transforms as inverses.
Lemma 1 (Möbius inversion: \(\lambda\) from \(y\)). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported, and let \(y : \mathbb{N}^k \to \mathbb{R}\) satisfy \(y = y_\lambda\) (Definition 1). Then for every \(\mathbf{d} \in \mathbb{N}^k\), \[\lambda_y(d_1, \dotsc, d_k) \;=\; \begin{cases} \lambda(d_1, \dotsc, d_k), & \text{if every $d_i$ is squarefree},\\ 0, & \text{otherwise}. \end{cases}\]
Proof. If some \(d_i\) is not squarefree, then every \(\mathbf{r}\) contributing to Definition 1 has \(d_i \mid r_i\) with \(r_i\) squarefree, forcing \(d_i\) squarefree — a contradiction. So the sum is empty and \(\lambda_y(\mathbf{d}) = 0\). Assume henceforth that every \(d_i\) is squarefree.
Both transforms are \(\mathbb{R}\)-linear in their argument, and every finitely supported function is a finite linear combination of point masses; so it suffices to prove the identity when \(\lambda = c \cdot \mathbf{1}_{\mathbf{d}^{(0)}}\) is a point mass at some \(\mathbf{d}^{(0)}\) with weight \(c\). For such \(\lambda\), \[y(\mathbf{r}) \;=\; \Bigl(\prod_i \mu(r_i)\varphi(r_i)\Bigr) \frac{c}{\prod_i d^{(0)}_i} \quad\text{if $r_i \mid d^{(0)}_i$ and $d^{(0)}_i$ is squarefree for all $i$,}\] and \(y(\mathbf{r}) = 0\) otherwise; in particular \(y\) is supported on the coordinatewise divisors of \(\mathbf{d}^{(0)}\). If some \(d^{(0)}_i\) fails to be squarefree, or if \(d_i \nmid d^{(0)}_i\) for some \(i\), then no \(\mathbf{r}\) contributes to \(\lambda_y(\mathbf{d})\) and both sides vanish (on the right, \(\lambda(\mathbf{d}) = 0\) unless \(\mathbf{d} = \mathbf{d}^{(0)}\), and the excluded cases exclude that). Otherwise substitute the displayed formula for \(y\) into Definition 1: the factors \(\varphi(r_i)\) cancel, and \[\lambda_y(\mathbf{d}) \;=\; \frac{c\,\prod_i \mu(d_i)\, d_i}{\prod_i d^{(0)}_i} \sum_{\substack{\mathbf{r}\\ d_i \mid r_i \mid d^{(0)}_i\ \forall i}} \prod_{i=1}^k \mu(r_i) \;=\; \frac{c\,\prod_i \mu(d_i)\, d_i}{\prod_i d^{(0)}_i} \prod_{i=1}^{k} \Bigl(\sum_{d_i \mid r_i \mid d^{(0)}_i} \mu(r_i)\Bigr),\] the sum having factorised across coordinates. For squarefree \(d^{(0)}_i\) the inner sum is \(\mu(d_i)\,\mathbf{1}[d_i = d^{(0)}_i]\), by Möbius inversion applied to the divisor chain \(d_i \mid r_i \mid d^{(0)}_i\) (writing \(r_i = d_i s_i\) with \(s_i \mid d^{(0)}_i/d_i\) and using \(\mu(d_i s_i) = \mu(d_i)\mu(s_i)\), valid since \(d^{(0)}_i\) is squarefree, together with \(\sum_{s \mid t} \mu(s) = \mathbf{1}[t = 1]\)). Hence the whole expression vanishes unless \(\mathbf{d} = \mathbf{d}^{(0)}\), in which case it equals \[\frac{c\,\prod_i \mu(d_i)\, d_i}{\prod_i d_i} \prod_i \mu(d_i) \;=\; c \prod_i \mu(d_i)^2 \;=\; c,\] because each \(d_i\) is squarefree. This is \(\lambda(\mathbf{d})\), as required. ◻
Lemma 1 (Möbius inversion: \(y\) from \(\lambda\)). Let \(y : \mathbb{N}^k \to \mathbb{R}\) be finitely supported and let \(\lambda_y\) be as in Definition 1. Then for every \(\mathbf{r} \in \mathbb{N}^k\), \[y_{\lambda_y}(r_1, \dotsc, r_k) \;=\; \begin{cases} y(r_1, \dotsc, r_k), & \text{if every $r_i$ is squarefree},\\ 0, & \text{otherwise}. \end{cases}\]
Proof. The argument is that of Lemma 1 with the roles of the two transforms exchanged. If some \(r_i\) is not squarefree, every \(\mathbf{d}\) contributing to Definition 1 at \(\mathbf{r}\) has \(r_i \mid d_i\) with \(d_i\) squarefree, which is impossible; so \(y_{\lambda_y}(\mathbf{r}) = 0\). Assume every \(r_i\) squarefree. By linearity reduce to \(y = c\cdot\mathbf{1}_{\mathbf{r}^{(0)}}\) a point mass. Then \(\lambda_y\) is supported on the coordinatewise divisors of \(\mathbf{r}^{(0)}\), and substituting Definition 1 into Definition 1 the factors \(d_i\) cancel against \(1/\prod_i d_i\), leaving \[y_{\lambda_y}(\mathbf{r}) \;=\; \frac{c\,\prod_i \mu(r_i)\varphi(r_i)}{\prod_i \varphi(r^{(0)}_i)} \prod_{i=1}^{k}\Bigl(\sum_{r_i \mid d_i \mid r^{(0)}_i} \mu(d_i)\Bigr).\] As before each inner sum equals \(\mu(r_i)\,\mathbf{1}[r_i = r^{(0)}_i]\), so the expression vanishes unless \(\mathbf{r} = \mathbf{r}^{(0)}\), and then equals \(c\prod_i \mu(r_i)^2 = c = y(\mathbf{r})\). ◻
Proof uses: lem_lambda_from_y
Together, Lemmas 1 and 1 say that the two transforms are mutually inverse \(\mathbb{R}\)-linear isomorphisms of the space of finitely supported functions \(\mathbb{N}^k \to \mathbb{R}\) that vanish at every tuple with some non-squarefree entry. In particular one may specify a sieve either by giving \(\lambda\) or by giving \(y\); the second is the form used in the variational problem.
The analysis of \(S_2^{(m)}\) requires a second, closely related transform, in which the \(m\)-th coordinate is frozen at \(1\) and the totient \(\varphi\) is replaced by the function \(g\) defined next.
On squarefree arguments — the only ones at which \(g\) is ever evaluated below, since it always occurs multiplied by \(\mu\) — one has \(g(n) = \prod_{p \mid n}(p - 2)\).
Definition 1 (The \(y^{(m)}\)-variables). Let \(m \in \{1, \dotsc, k\}\) and let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be finitely supported. Define \(y^{(m)} : \mathbb{N}^k \to \mathbb{R}\) by \[y^{(m)}(r_1, \dotsc, r_k) \;:=\; \Bigl(\prod_{i=1}^{k} \mu(r_i)\, g(r_i)\Bigr) \sum_{\substack{\mathbf{d} \in \mathbb{N}^k,\ d_m = 1 \\ r_i \mid d_i \text{ and } d_i \text{ squarefree } \forall i}} \frac{\lambda(d_1, \dotsc, d_k)}{\prod_{i=1}^{k} \varphi(d_i)}.\]
Uses: def_mobius, def_g_func, def_totient, not_bold_tuple
Because \(\mu(1) = g(1) = \varphi(1) = 1\), the factors carrying the index \(m\) contribute trivially: the prefactor may equally be written \(\prod_{i \ne m} \mu(r_i)g(r_i)\) and the denominator \(\prod_{i \ne m}\varphi(d_i)\), which is the shape in which \(y^{(m)}\) appears in the evaluation of \(S_2^{(m)}\).
Lemma 1 (Support of \(y^{(m)}\)). Let \(m \in \{1, \dotsc, k\}\) and \(\lambda : \mathbb{N}^k \to \mathbb{R}\) finitely supported. If \(y^{(m)}(r_1, \dotsc, r_k) \ne 0\), then \(r_m = 1\).
Uses: def_ym_vars
Proof. If \(y^{(m)}(\mathbf{r}) \ne 0\) then the defining sum of Definition 1 has a nonzero summand, so there is a \(\mathbf{d}\) with \(d_m = 1\) and \(r_i \mid d_i\) for every \(i\). Taking \(i = m\) gives \(r_m \mid 1\), hence \(r_m = 1\). ◻
Finally we record the coefficient function that the two instantiations use in \(S_1\) and \(S_2\). It is specified in the \(y\)-variables — a rescaled sample of a fixed function \(F\) on the truncation scale — and the corresponding \(\lambda\) is produced by the Möbius transform of Definition 1. The analytic input is \(F\), and \(\lambda\) is the derived object.
By construction \(y_0\) has permissible support at level \(T\), hence is finitely supported, so Definition 1 applies to it.
Since the Möbius transform preserves permissible support, \(\lambda_0\) has permissible support at level \(T\), so all of Definitions 1–1 apply to it. Instantiation I takes \(T = \lfloor R\rfloor\) with \(F\) supported in the standard simplex \(\mathcal{R}_k\); Instantiation II takes \(T = \lfloor R^{1+\varepsilon}\rfloor\) with \(F\) supported in the enlarged simplex \(\mathcal{T}_\varepsilon\). Everything above is common to both.
Definition 1 (Maximal modulus of a finitely supported system). Let \(A\) be a set and let \(f : A \to \mathbb{R}\) be finitely supported. Its maximal modulus is \[\|f\|_{\max} \;:=\; \max_{a \in \operatorname{supp} f} |f(a)|,\] with the convention \(\|f\|_{\max} := 0\) when \(f\) vanishes identically.
Definition 1 (The truncated weight sum \(H(z)\)). Let \(S\) be a sieve datum with associated weight \(h\) (Definition 1). For a real \(z > 0\) set \[H(z) \;:=\; \sum_{0 < d < z} h(d),\] the sum taken over positive integers \(d\) strictly below \(z\).
Uses: def_h_conv
Definition 1 (The Maynard sieve datum). Let \(N\) be a positive integer with \(2 \le D_0(N)\). The Maynard sieve datum at \(N\) is the sieve datum (Definition 1) whose density is the \(W\)-tricked density \(\gamma_g\) of Definition 1, whose modulus is \(V := W\), and whose growth and structure constants are \(A_1 := 1/2\) and \(A_3 := 2\). Its remaining constants \(A_2\) and \(L\) are the deviation bounds carried by the datum; that they may be taken uniformly in \(N\) is established in Section 6.
Uses: def_sieve_datum, def_W_trick, def_gamma_g
The variational problem
The sieve estimates of the preceding sections reduce the problem to one in the calculus of variations. A profile \(F\) supported on a simplex determines the sieve coefficients \(\lambda\); the leading term of \(S_1(\lambda)\) is the \(L^2\) mass \(I_k(F)\) of the profile, and the leading term of \(S_2^{(m)}(\lambda)\) is the \(L^2\) mass \(J_{k,\varepsilon}^{(m)}(F)\) of its \(m\)-th marginal. The supremum of the ratio \(\bigl(\sum_m J_{k,\varepsilon}^{(m)}(F)\bigr)/I_k(F)\) over profiles \(F\) is the constant \(M_{k,\varepsilon}\).
The construction is carried out for an enlargement parameter \(\varepsilon \ge 0\). The profile is supported on the enlarged simplex \(\mathcal{T}_\varepsilon\) of total mass \(1 + \varepsilon\), while each marginal is taken over the shrunken slice \(\mathcal{S}_\varepsilon\) of total mass \(1 - \varepsilon\); at \(\varepsilon = 0\) both reduce to the standard simplex and its faces, and every statement below specialises to the classical one. The denominator \(I_k(F) = \int F^2\) does not depend on \(\varepsilon\); the enlargement enters only through the wider class of admissible competitors.
We introduce \(\mathcal{T}_\varepsilon\), \(F\), \(I_k\), \(J_{k,\varepsilon}^{(m)}\), \(M_{k,\varepsilon}\) and \(r_k\), record the basic properties of \(M_{k,\varepsilon}\) (finiteness, positivity, \(L^2\)-continuity of the two functionals, density of smooth profiles), and prove the variational criterion (Lemma 1), which passes from a lower bound on \(M_{k,\varepsilon}\) to the strict inequality used by the sieve estimates. Both cases invoke Lemma 1, one at \(\varepsilon = 0\), the other at \(\varepsilon = 1/25\). The remaining lemmas bound \(M_{k,\varepsilon}\) from below by evaluating the two functionals on explicit polynomial profiles.
The simplex and the profile
Notation 1 (Power sums). For \(j \in \mathbb{Z}_{\ge 1}\) we write \(P_j\) for the \(j\)-th power sum in the coordinates \(t_1, \dotsc, t_k\), \[P_j(t_1, \dotsc, t_k) \;:=\; \sum_{i=1}^{k} t_i^{\,j} ,\] so that \(P_1 = t_1 + \dotsb + t_k\) and \(P_2 = t_1^2 + \dotsb + t_k^2\). When one coordinate is distinguished we write \(P_j'\) for the corresponding power sum in the remaining \(k - 1\) coordinates.
Definition 1 (Scaled standard simplex). Let \(k \in \mathbb{Z}_{\ge 0}\) and \(s \in \mathbb{R}\). The scaled standard simplex of dimension \(k\) and scale \(s\) is \[\mathcal{R}_k(s) \;:=\; \Bigl\{\, t = (t_1, \dotsc, t_k) \in \mathbb{R}^k \;:\; t_i \ge 0 \text{ for every } i, \ \sum_{i=1}^{k} t_i \le s \,\Bigr\}.\] Three instances recur, and we abbreviate them. For \(\varepsilon \ge 0\): \[\mathcal{R}_k \;:=\; \mathcal{R}_k(1), \qquad \mathcal{T}_\varepsilon \;:=\; \mathcal{R}_k(1 + \varepsilon), \qquad \mathcal{S}_\varepsilon \;:=\; \mathcal{R}_{k-1}(1 - \varepsilon) ,\] called respectively the standard simplex, the \(\varepsilon\)-enlarged simplex and the \(\varepsilon\)-shrunken slice. Thus \(\mathcal{T}_0 = \mathcal{R}_k\) and \(\mathcal{S}_0 = \mathcal{R}_{k-1}\); we also write \(\mathcal{R}_k^{(m)}\) for the copy of \(\mathcal{R}_{k-1}\) obtained by deleting the \(m\)-th coordinate of \(\mathbb{R}^k\), so that \(\mathcal{S}_0 = \mathcal{R}_k^{(m)}\) under that identification.
Definition 1 (Maynard’s combinatorial polynomial \(G_{b,j}\)). For \(b, j, x \in \mathbb{Z}_{\ge 0}\) set \[G_{b,j}(x) \;:=\; b! \sum_{\substack{(d_1, \dotsc, d_x) \in \mathbb{Z}_{\ge 0}^{x} \\ d_1 + \dotsb + d_x = b}} \ \prod_{i=1}^{x} \frac{(j\, d_i)!}{d_i!} ,\] the sum being over all \(x\)-tuples of non-negative integers with total \(b\). Each quotient \((j d_i)!/d_i!\) is an integer, so \(G_{b,j}(x) \in \mathbb{Z}_{\ge 0}\). Discarding the zero entries of a tuple and recording only the \(r\) positive ones — each reduced tuple being counted \(\binom{x}{r}\) times, since the weight takes the value \(1\) at \(d_i = 0\) — gives the equivalent form \[G_{b,j}(x) \;=\; b! \sum_{r=1}^{b} \binom{x}{r} \sum_{\substack{d_1 + \dotsb + d_r = b \\ d_i \ge 1}} \prod_{i=1}^{r} \frac{(j d_i)!}{d_i!} \qquad (b \ge 1),\] in which \(x\) occurs only inside the binomial coefficients; this exhibits \(G_{b,j}\) as a polynomial in \(x\) of degree \(b\).
Definition 1 (Extension by zero from the simplex). Let \(k \in \mathbb{N}\) and let \(P : \mathbb{R}^k \to \mathbb{R}\). The profile attached to \(P\) is \[F_P \;:=\; P \cdot \mathbf{1}_{\mathcal{R}_k}, \qquad\text{that is,}\qquad F_P(t) = \begin{cases} P(t), & t \in \mathcal{R}_k,\\ 0, & t \notin \mathcal{R}_k.\end{cases}\]
Uses: def_simplex
Definition 1 (The monomial index set \(\mathcal{B}_d\)). For \(d \in \mathbb{Z}_{\ge 0}\) set \[\mathcal{B}_d \;:=\; \bigl\{\, (b, c) \in \mathbb{Z}_{\ge 0}^2 \;:\; b + 2c \le d \,\bigr\}.\] A pair \((b, c) \in \mathcal{B}_d\) indexes the symmetric polynomial \((1 - P_1)^b\, P_2^{\,c}\) in \(t_1, \dotsc, t_k\); these polynomials span the space of profiles used to bound \(M_{k,0}\) from below.
Uses: not_power_sum
The functionals \(I_k\), \(J_{k,\varepsilon}^{(m)}\) and the constant \(M_{k,\varepsilon}\)
Definition 1 (The functional \(I_k\)). For \(k \in \mathbb{N}\) and \(F \in L^2(\mathbb{R}^k)\), \[I_k(F) \;:=\; \int_{\mathbb{R}^k} F(t_1, \dotsc, t_k)^2 \, dt_1 \dotsm dt_k .\] If \(F\) vanishes outside a measurable set \(S\) then \(I_k(F) = \int_S F^2 = \|F\|_{L^2(S)}^2\); in particular \(I_k(F) = \|F\|^2_{L^2(\mathcal{T}_\varepsilon)}\) for a profile supported in \(\mathcal{T}_\varepsilon\). The functional carries no dependence on \(\varepsilon\).
Uses: def_simplex
Definition 1 (The marginal functional \(J_{k,\varepsilon}^{(m)}\)). Let \(k \ge 1\), let \(\varepsilon \ge 0\), let \(m \in \{1, \dotsc, k\}\), and let \(F : \mathbb{R}^k \to \mathbb{R}\) vanish outside \(\mathcal{T}_\varepsilon\). Write \(T_m\) for the \(m\)-th marginal operator \[(T_m F)(t_1, \dotsc, \widehat{t_m}, \dotsc, t_k) \;:=\; \int_0^{\infty} F(t_1, \dotsc, t_k)\, dt_m ,\] defined for almost every \((t_i)_{i \ne m} \in \mathbb{R}^{k-1}\), and set \[J_{k,\varepsilon}^{(m)}(F) \;:=\; \int_{\mathcal{S}_\varepsilon} (T_m F)^2 \, \prod_{i \ne m} dt_i \;=\; \|T_m F\|^2_{L^2(\mathcal{S}_\varepsilon)} .\] We abbreviate \(J_k^{(m)} := J_{k,0}^{(m)}\), the marginal functional of the standard problem, whose outer domain \(\mathcal{S}_0\) is \(\mathcal{R}_k^{(m)}\).
Uses: def_simplex
Definition 1 (The variational constant \(M_{k,\varepsilon}\)). For \(k \ge 1\) and \(\varepsilon \ge 0\), \[M_{k,\varepsilon} \;:=\; \sup\left\{\ \frac{\sum_{m=1}^{k} J_{k,\varepsilon}^{(m)}(F)}{I_k(F)} \ \middle|\ F \in L^2(\mathcal{T}_\varepsilon),\ I_k(F) \ne 0 \ \right\} .\] The supremum is taken over the whole of \(L^2(\mathcal{T}_\varepsilon)\), each competitor being extended by zero outside \(\mathcal{T}_\varepsilon\): no smoothness, no symmetry and no sign condition is imposed on \(F\). We abbreviate \(M_k := M_{k,0}\).
Definition 1 (The prime count \(r_k(\vartheta)\)). For \(\vartheta \in (0,1)\) and \(k \ge 1\), \[r_k(\vartheta) \;:=\; \Bigl\lceil \frac{\vartheta\, M_k}{2} \Bigr\rceil \in \mathbb{Z}_{\ge 0},\] where \(M_k = M_{k,0}\). This is the number of primes the sieve delivers among \(n + h_1, \dotsc, n + h_k\) at level of distribution \(\vartheta\).
Uses: def_M_k, def_level_of_distribution
From a profile to the sieve coefficients
Definition 1 (The \(y\)-variables attached to a profile). Let \(k \in \mathbb{N}\), let \(R > 1\) and \(W \in \mathbb{N}\), and let \(F : \mathbb{R}^k \to \mathbb{R}\). Define \(y_0 : \mathbb{N}^k \to \mathbb{R}\) by \[y_0(r_1, \dotsc, r_k) \;:=\; \begin{cases} F\!\left(\dfrac{\log r_1}{\log R}, \dotsc, \dfrac{\log r_k}{\log R}\right), & (r_1, \dotsc, r_k) \text{ admissible for } (R, W),\\[1.2ex] 0, & \text{otherwise}, \end{cases}\] where a tuple is admissible for \((R, W)\) in the sense of Definition 1, i.e. \(\prod_i r_i \le R\), \(\gcd(\prod_i r_i, W) = 1\) and \(\prod_i r_i\) is squarefree.
Uses: def_adm_support
Definition 1 (The sieve coefficients attached to a profile). With \(k\), \(R\), \(W\) and \(F\) as in Definition 1, the coefficients \(\lambda_0 : \mathbb{N}^k \to \mathbb{R}\) attached to \(F\) are the image of \(y_0\) under the \(y \mapsto \lambda\) correspondence of Definition 1: \[\lambda_0(d_1, \dotsc, d_k) \;=\; \Bigl(\prod_{i=1}^{k} \mu(d_i)\, d_i\Bigr) \sum_{\substack{(r_1, \dotsc, r_k) \text{ admissible for } (R, W) \\ d_i \mid r_i \text{ for every } i}} \frac{y_0(r_1, \dotsc, r_k)}{\prod_{i=1}^{k} \varphi(r_i)} .\] Taking \(F\) supported in \(\mathcal{T}_\varepsilon\) and truncation level \(R^{1+\varepsilon}\) produces the enlarged weight; taking \(F\) supported in \(\mathcal{R}_k\) and level \(R\) produces the standard one.
Basic properties of \(I_k\), \(J_{k,\varepsilon}^{(m)}\) and \(M_{k,\varepsilon}\)
Lemma 1 (Each marginal is dominated by the \(L^2\) mass). Let \(k \ge 1\), let \(\varepsilon \ge 0\), let \(F \in L^2(\mathcal{T}_\varepsilon)\) and let \(m \in \{1, \dotsc, k\}\). Then \[J_{k,\varepsilon}^{(m)}(F) \;\le\; (1 + \varepsilon)\, I_k(F).\]
Uses: def_I_k, def_J_k, def_simplex
Proof. Fix \((t_i)_{i \ne m} \in \mathcal{S}_\varepsilon\) and consider the fibre \[E \;:=\; \{\, s \ge 0 : (t_1, \dotsc, t_{m-1}, s, t_{m+1}, \dotsc, t_k) \in \mathcal{T}_\varepsilon \,\} .\] Since every \(t_i \ge 0\), the fibre is the interval \([0,\, 1 + \varepsilon - \sum_{i \ne m} t_i]\) (empty if the right endpoint is negative), whose length is at most \(1 + \varepsilon\). Since \(F\) vanishes off \(\mathcal{T}_\varepsilon\), the Cauchy–Schwarz inequality on \(E\) gives \[\bigl|(T_m F)\bigl((t_i)_{i \ne m}\bigr)\bigr|^2 \;=\; \Bigl|\int_E F \, ds\Bigr|^2 \;\le\; |E| \int_E F^2 \, ds \;\le\; (1 + \varepsilon) \int_E F^2 \, ds .\] Integrating over \((t_i)_{i \ne m} \in \mathcal{S}_\varepsilon\) and applying Tonelli’s theorem to the right-hand side, the iterated integral is bounded by the integral of \(F^2\) over all of \(\mathbb{R}^k\), which is \(I_k(F)\). Hence \(J_{k,\varepsilon}^{(m)}(F) \le (1 + \varepsilon) I_k(F)\).
Equivalently: \(T_m\) is a bounded linear operator \(L^2(\mathcal{T}_\varepsilon) \to L^2(\mathcal{S}_\varepsilon)\) whose norm is at most the square root of the essential supremum of the fibre lengths, and that essential supremum is at most \(1 + \varepsilon\). ◻
Lemma 1 (Iterated form of the standard marginal functional). Let \(k \ge 1\), let \(m \in \{1, \dotsc, k\}\), and let \(F \in L^2(\mathcal{R}_k)\) vanish outside \(\mathcal{R}_k\). Then \[J_k^{(m)}(F) \;=\; \int_{\mathcal{R}_k^{(m)}} \Bigl( \int_0^{\,1 - \sum_{i \ne m} t_i} F(t_1, \dotsc, t_{m-1}, s, t_{m+1}, \dotsc, t_k)\, ds \Bigr)^{\!2} \prod_{i \ne m} dt_i .\]
Uses: def_J_k, def_simplex
Proof. At \(\varepsilon = 0\) the outer domain \(\mathcal{S}_0\) is \(\mathcal{R}_k^{(m)}\), so only the inner integral has to be identified. Fix \((t_i)_{i \ne m} \in \mathcal{R}_k^{(m)}\). Then every \(t_i \ge 0\) and \(\sum_{i \ne m} t_i \le 1\), and the fibre of \(\mathcal{R}_k\) over this point is exactly \([0,\, 1 - \sum_{i \ne m} t_i]\). For \(s\) beyond that endpoint the inserted point lies outside \(\mathcal{R}_k\), so \(F\) vanishes there by hypothesis; hence the integral \(\int_0^{\infty} F\, dt_m\) defining \(T_m F\) agrees with the truncated integral \(\int_0^{1 - \sum_{i \ne m} t_i} F\, ds\). Squaring and integrating over \(\mathcal{R}_k^{(m)}\) gives the claim.
The hypothesis that \(F\) vanish pointwise outside \(\mathcal{R}_k\), and not merely almost everywhere on \(\mathcal{R}_k\), is essential: \(T_m\) samples \(F\) along the whole half-line, so a modification of \(F\) on a set that is null in \(\mathcal{R}_k\) but not in \(\mathbb{R}^k\) would change the left-hand side. ◻
Proof uses: lem_J_le_I
Lemma 1 (\(L^2\)-continuity of \(I_k\)). For every \(\varepsilon \ge 0\) the functional \(I_k : L^2(\mathcal{T}_\varepsilon) \to \mathbb{R}\) is continuous.
Uses: def_I_k, def_simplex
Proof. On \(L^2(\mathcal{T}_\varepsilon)\) we have \(I_k(F) = \|F\|_{L^2(\mathcal{T}_\varepsilon)}^2\) by Definition 1, so \(I_k\) is the square of the norm, and the norm on a normed space is continuous. Quantitatively, for \(F, G \in L^2(\mathcal{T}_\varepsilon)\) the factorisation \(F^2 - G^2 = (F - G)(F + G)\) together with Cauchy–Schwarz and the triangle inequality gives \[|I_k(F) - I_k(G)| \;\le\; \|F - G\|\,\|F + G\| \;\le\; \bigl(\|F\| + \|G\|\bigr)\,\|F - G\| ,\] all norms being those of \(L^2(\mathcal{T}_\varepsilon)\). ◻
Lemma 1 (\(L^2\)-continuity of \(J_{k,\varepsilon}^{(m)}\)). For every \(k \ge 1\), every \(\varepsilon \ge 0\) and every \(m \in \{1, \dotsc, k\}\), the functional \(J_{k,\varepsilon}^{(m)} : L^2(\mathcal{T}_\varepsilon) \to \mathbb{R}\) is continuous.
Uses: def_J_k, def_simplex
Proof. The marginal operator \(T_m\) is linear, and by the proof of Lemma 1 it is bounded from \(L^2(\mathcal{T}_\varepsilon)\) to \(L^2(\mathcal{S}_\varepsilon)\), with norm at most \((1 + \varepsilon)^{1/2}\); hence it is continuous. The map \(H \mapsto \|H\|^2_{L^2(\mathcal{S}_\varepsilon)}\) is continuous, being the square of a norm. By Definition 1, \(J_{k,\varepsilon}^{(m)}\) is the composite of these two continuous maps. Quantitatively, the argument of Lemma 1 applied in \(L^2(\mathcal{S}_\varepsilon)\) together with the operator bound gives \[\bigl|J_{k,\varepsilon}^{(m)}(F) - J_{k,\varepsilon}^{(m)}(G)\bigr| \;\le\; (1 + \varepsilon)\bigl(\|F\| + \|G\|\bigr)\,\|F - G\|\] with the norms of \(L^2(\mathcal{T}_\varepsilon)\). ◻
Proof uses: lem_J_le_I, lem_Ik_L2_continuous
Lemma 1 (Upper bound for \(M_{k,\varepsilon}\)). For every \(k \ge 1\) and every \(\varepsilon \ge 0\) the supremum defining \(M_{k,\varepsilon}\) is finite, and \[M_{k,\varepsilon} \;\le\; (1 + \varepsilon)\, k .\] In particular \(M_k \le k\).
Uses: def_M_k, def_simplex
Proof. Let \(F \in L^2(\mathcal{T}_\varepsilon)\) with \(I_k(F) \ne 0\). Summing the bound of Lemma 1 over the \(k\) values of \(m\) gives \(\sum_{m=1}^{k} J_{k,\varepsilon}^{(m)}(F) \le (1 + \varepsilon) k\, I_k(F)\), so the Rayleigh quotient at \(F\) is at most \((1+\varepsilon)k\). The family of quotients is therefore bounded above, its supremum exists in \(\mathbb{R}\), and it is at most \((1 + \varepsilon)k\). ◻
Proof uses: lem_J_le_I
Lemma 1 (Positivity of \(M_{k,\varepsilon}\)). For every \(k \ge 1\) and every \(\varepsilon \in [0, 1)\) we have \(M_{k,\varepsilon} > 0\).
Uses: def_M_k, def_simplex
Proof. Take \(F := \mathbf{1}_{\mathcal{R}_k}\), which vanishes outside \(\mathcal{R}_k \subseteq \mathcal{T}_\varepsilon\) and so is an element of \(L^2(\mathcal{T}_\varepsilon)\). The simplex \(\mathcal{R}_k\) contains the non-empty open set \(\{t : t_i > 0 \ \forall i,\ \sum_i t_i < 1\}\), hence has positive volume, and \(I_k(F) = \operatorname{vol} (\mathcal{R}_k) > 0\); so \(F\) is an admissible competitor.
Fix \(m\). For \((t_i)_{i \ne m}\) with \(t_i \ge 0\) and \(\sum_{i \ne m} t_i \le 1\) the fibre of \(\mathcal{R}_k\) over that point is the interval \([0,\, 1 - \sum_{i \ne m} t_i]\), so, exactly as in the proof of Lemma 1, \((T_m F)((t_i)_{i \ne m}) = 1 - \sum_{i \ne m} t_i\). The shrunken slice \(\mathcal{S}_\varepsilon\) is contained in \(\mathcal{R}_{k-1}\) because \(\varepsilon \ge 0\), and it contains the non-empty open set \(V := \{t : t_i > 0 \ \forall i,\ \sum_i t_i < (1 - \varepsilon)/2\}\) because \(\varepsilon < 1\); on \(V\) the integrand \((1 - \sum_{i \ne m} t_i)^2\) exceeds \(1/4\). Hence \[J_{k,\varepsilon}^{(m)}(F) \;\ge\; \tfrac14 \operatorname{vol}(V) \;>\; 0 .\] Therefore \(\sum_m J_{k,\varepsilon}^{(m)}(F) > 0\), the Rayleigh quotient at \(F\) is strictly positive, and \(M_{k,\varepsilon}\), being at least that quotient, is strictly positive. ◻
Proof uses: lem_J_iterated
Smooth profiles and the variational criterion
The sieve asymptotics for \(S_1\) and \(S_2^{(m)}\) require a smooth profile supported in \(\mathcal{T}_\varepsilon\), whereas \(M_{k,\varepsilon}\) is a supremum over all of \(L^2(\mathcal{T}_\varepsilon)\). Smooth functions supported in the simplex are \(L^2\)-dense there, hence (by the continuity of \(I_k\) and \(J_{k,\varepsilon}^{(m)}\)) they realise the supremum to within any prescribed error, and the resulting near-maximiser satisfies the strict inequality required by the sieve estimates. The density statement is proved on the standard simplex and transported to \(\mathcal{T}_\varepsilon\) by the dilation \(t \mapsto (1 + \varepsilon)^{-1} t\), which carries \(\mathcal{T}_\varepsilon\) onto \(\mathcal{R}_k\).
Lemma 1 (Smooth cut-off to an inset of \(\mathcal{R}_k\)). Let \(k \ge 2\) and let \(\eta\) satisfy \(0 < \eta < 1/(2k)\). There exists a smooth function \(\chi : \mathbb{R}^k \to [0,1]\) such that
\(\chi(t) = 1\) for every \(t\in\mathcal{R}_k\) satisfying \(t_i\ge\eta\) for all \(i\) and \(\sum_i t_i\le 1-\eta\);
\(\chi(t) = 0\) for every \(t \notin \mathcal{R}_k\);
for every measurable \(G : \mathbb{R}^k \to \mathbb{R}\) and every \(M \ge 0\) with \(|G| \le M\) on \(\mathcal{R}_k\), \[\int_{\mathcal{R}_k} \bigl(G - G\chi\bigr)^2 \;\le\; M^2 \operatorname{vol}\bigl(\{t\in\mathcal{R}_k:t_i<\eta\text{ for some }i \text{ or }\sum_i t_i>1-\eta\}\bigr) \;\le\; M^2 \cdot 2k\eta .\]
Uses: def_simplex
Proof. Fix a smooth non-decreasing \(\psi : \mathbb{R} \to [0,1]\) with \(\psi(u) = 0\) for \(u \le 0\) and \(\psi(u) = 1\) for \(u \ge 1\), and put \(\psi_\eta(u) := \psi(u/\eta)\). Set \[\chi(t_1, \dotsc, t_k) \;:=\; \Bigl(\prod_{i=1}^{k} \psi_\eta(t_i)\Bigr) \cdot \psi_\eta\Bigl(1 - \sum_{i=1}^{k} t_i\Bigr).\] This is a finite product of smooth functions, hence smooth, with values in \([0,1]\).
(i) If \(t \in \mathcal{R}_k^{(\eta)}\) then \(t_i \ge \eta\) for each \(i\) and \(\sum_i t_i \le 1 - \eta\), so every factor equals \(1\).
(ii) If \(t \notin \mathcal{R}_k\) then either \(t_i < 0\) for some \(i\), or \(\sum_i t_i > 1\); in the first case \(\psi_\eta(t_i) = 0\) and in the second \(\psi_\eta(1 - \sum_i t_i) = 0\).
(iii) On \(\mathcal{R}_k^{(\eta)}\) we have \(G - G\chi = 0\) by (i), while on all of \(\mathcal{R}_k\) we have \(|G - G\chi| = |G|\,|1 - \chi| \le M\). Hence \[\int_{\mathcal{R}_k} (G - G\chi)^2 = \int_{\mathcal{R}_k \setminus \mathcal{R}_k^{(\eta)}} (G - G\chi)^2 \le M^2 \operatorname{vol}\bigl(\mathcal{R}_k \setminus \mathcal{R}_k^{(\eta)}\bigr).\] For the second inequality, \(\mathcal{R}_k \setminus \mathcal{R}_k^{(\eta)}\) is covered by the \(k\) slabs \(\{t \in \mathcal{R}_k : t_i < \eta\}\) together with the slab \(\{t \in \mathcal{R}_k : \sum_i t_i > 1 - \eta\}\). Each of the first \(k\) slabs has volume at most \(\eta/(k-1)!\), since its cross-section at fixed \(t_i\) lies in a \((k-1)\)-dimensional simplex of volume \(1/(k-1)!\); the last slab, being the difference of two simplices, has volume \(\bigl(1 - (1-\eta)^k\bigr)/k! \le \eta/(k-1)!\). Summing gives \(\operatorname{vol}(\mathcal{R}_k \setminus \mathcal{R}_k^{(\eta)}) \le (k+1)\eta/(k-1)! \le 2k\eta\) for \(k \ge 2\). ◻
Lemma 1 (\(L^2\)-density of smooth functions supported in \(\mathcal{R}_k\)). Let \(k \ge 2\), let \(F_0 \in L^2(\mathcal{R}_k)\) and let \(\delta > 0\). Then there exists a smooth \(F_1 : \mathbb{R}^k \to \mathbb{R}\) whose closed support is contained in \(\mathcal{R}_k\) and which satisfies \[\|F_0 - F_1\|_{L^2(\mathcal{R}_k)} \;<\; \delta .\]
Uses: def_simplex
Proof. Extend \(F_0\) by zero to \(\mathbb{R}^k\); the extension lies in \(L^2(\mathbb{R}^k)\) with the same norm. By the density of smooth compactly supported functions in \(L^2(\mathbb{R}^k)\) there is a smooth compactly supported \(\widetilde F\) with \(\|F_0 - \widetilde F\|_{L^2(\mathbb{R}^k)} < \delta/3\); restricting the integration to \(\mathcal{R}_k\) only decreases the left-hand side. Being continuous with compact support, \(\widetilde F\) is bounded on \(\mathcal{R}_k\); let \(M \ge 0\) be a bound for \(|\widetilde F|\) there.
Choose \(\eta \in (0, 1/(2k))\) so small that \(M\sqrt{2k\eta} < \delta/3\), and let \(\chi\) be the cut-off supplied by Lemma 1 for this \(\eta\). Put \(F_1 := \widetilde F \cdot \chi\). Then \(F_1\) is smooth, and its closed support is contained in that of \(\chi\), hence in the closed set \(\mathcal{R}_k\). By part (iii) of Lemma 1 applied with \(G = \widetilde F\), \[\|\widetilde F - F_1\|_{L^2(\mathcal{R}_k)} \;\le\; M\sqrt{2k\eta} \;<\; \delta/3 ,\] and the triangle inequality gives \(\|F_0 - F_1\|_{L^2(\mathcal{R}_k)} < 2\delta/3 < \delta\). ◻
Proof uses: lem_smooth_cutoff_simplex
Lemma 1 (Smooth approximation of the variational quotient). Let \(k \ge 2\), let \(f \in L^2(\mathcal{R}_k)\) with \(I_k(f) > 0\), and let \(\eta > 0\). Then there is \(F \in C^\infty(\mathbb{R}^k)\) with \(\operatorname{supp} F \subset \mathcal{R}_k\), \(I_k(F) > 0\) and \[\Bigl| \frac{\sum_{m=1}^k J_k^{(m)}(F)}{I_k(F)} \;-\; \frac{\sum_{m=1}^k J_k^{(m)}(f)}{I_k(f)} \Bigr| \;<\; \eta .\]
Uses: def_simplex, def_I_k, def_J_k
Proof. Write \(\mathcal{Q}(g) := \bigl(\sum_m J_k^{(m)}(g)\bigr)/I_k(g)\) for the variational quotient, defined on the set where \(I_k \ne 0\). Its numerator is a finite sum of the functionals \(J_k^{(m)}\), continuous by Lemma 1, and its denominator \(I_k\) is continuous by Lemma 1 and is non-zero at \(f\); hence \(\mathcal{Q}\) is continuous at \(f\).
Choose \(\delta_1 > 0\) such that \(\|g - f\|_{L^2(\mathcal{R}_k)} < \delta_1\) implies \(|\mathcal{Q}(g) - \mathcal{Q}(f)| < \eta\), and, using continuity of \(I_k\) at \(f\) together with \(I_k(f) > 0\), choose \(\delta_2 > 0\) such that \(\|g - f\|_{L^2(\mathcal{R}_k)} < \delta_2\) implies \(I_k(g) > 0\). Apply Lemma 1 with \(\delta := \min(\delta_1, \delta_2)\) to obtain a smooth \(F\) supported on \(\mathcal{R}_k\) with \(\|f - F\|_{L^2(\mathcal{R}_k)} < \delta\). Both conclusions hold for this \(F\). ◻
The main inequality on \(S\)
Lemma 1 (Smooth near-maximisers for \(M_{k,\varepsilon}\)). Let \(k \ge 2\), let \(\varepsilon \in [0,1)\) and let \(\delta > 0\). Then there exists a smooth \(F : \mathbb{R}^k \to \mathbb{R}\) whose closed support is contained in \(\mathcal{T}_\varepsilon\), with \(I_k(F) > 0\) and \[(M_{k,\varepsilon} - \delta)\, I_k(F) \;<\; \sum_{m=1}^{k} J_{k,\varepsilon}^{(m)}(F).\]
Uses: def_M_k, def_I_k, def_J_k, def_simplex
Proof. Step 1: smooth functions supported in \(\mathcal{T}_\varepsilon\) are \(L^2\)-dense there. Let \(\sigma := 1 + \varepsilon > 0\) and let \(\Lambda_\varepsilon\) be the dilation \[(\Lambda_\varepsilon G)(t) \;:=\; \sigma^{-k/2}\, G(\sigma^{-1} t) .\] Since \(\sigma^{-1}\mathcal{T}_\varepsilon = \mathcal{R}_k\), the map \(\Lambda_\varepsilon\) carries functions supported in \(\mathcal{R}_k\) to functions supported in \(\mathcal{T}_\varepsilon\), preserves smoothness, and is an isometry of \(L^2(\mathcal{R}_k)\) onto \(L^2(\mathcal{T}_\varepsilon)\) by the change of variables \(t = \sigma u\). Composing Lemma 1 with \(\Lambda_\varepsilon\) shows: for every \(H_0 \in L^2(\mathcal{T}_\varepsilon)\) and every \(\delta' > 0\) there is a smooth \(H_1\) with closed support in \(\mathcal{T}_\varepsilon\) and \(\|H_0 - H_1\|_{L^2(\mathcal{T}_\varepsilon)} < \delta'\).
Step 2: transfer along the Rayleigh quotient. By Lemma 1 we have \(M_{k,\varepsilon} > 0\), and by Lemma 1 the supremum is finite. Write \(Q(G) := \bigl(\sum_m J_{k,\varepsilon}^{(m)}(G)\bigr)/I_k(G)\). Since \(M_{k,\varepsilon}\) is the supremum of \(Q\), there is \(G_0 \in L^2(\mathcal{T}_\varepsilon)\) with \[\max\Bigl(M_{k,\varepsilon} - \tfrac{\delta}{2},\ \tfrac{M_{k,\varepsilon}}{2}\Bigr) \;<\; Q(G_0),\] and in particular \(I_k(G_0) > 0\), since a competitor must have non-zero mass. Both \(\sum_m J_{k,\varepsilon}^{(m)}\) and \(I_k\) are continuous on \(L^2(\mathcal{T}_\varepsilon)\) (Lemmas 1 and 1), so \(Q\) is continuous at \(G_0\) and \(I_k\) stays positive near \(G_0\): there is \(\delta' > 0\) such that every \(H\) with \(\|G_0 - H\| < \delta'\) satisfies \(I_k(H) > 0\) and \(|Q(H) - Q(G_0)| < \delta/2\).
By Step 1 choose such an \(H = F\) smooth with closed support in \(\mathcal{T}_\varepsilon\). Then \(I_k(F) > 0\) and \(Q(F) > Q(G_0) - \delta/2 > M_{k,\varepsilon} - \delta\); multiplying by \(I_k(F) > 0\) gives the asserted inequality. ◻
The following lemma is used by both cases of the argument. It converts a lower bound on \(M_{k,\varepsilon}\) into the strict inequality \[\rho\, I_k(F) \;<\; \Bigl(\frac{\vartheta}{2} - \delta\Bigr) \sum_{m=1}^{k} J_{k,\varepsilon}^{(m)}(F)\] for a smooth profile \(F\) — exactly the hypothesis under which the \(S_1\) and \(S_2^{(m)}\) asymptotics force \(S_2(\lambda_0) - \rho\, S_1(\lambda_0) > 0\) for all large \(N\). Taking \(\varepsilon = 0\) gives the standard criterion; taking \(\varepsilon > 0\) gives the enlarged one.
Lemma 1 (The variational criterion). Let \(k \ge 2\), let \(\varepsilon \in [0, 1)\), let \(\vartheta \in (0,1)\) and let \(\rho > 0\). Suppose that \[M_{k,\varepsilon} \;>\; \frac{2\rho}{\vartheta} .\] Then there exist a smooth \(F : \mathbb{R}^k \to \mathbb{R}\) whose closed support is contained in \(\mathcal{T}_\varepsilon\) and which satisfies \(I_k(F) > 0\), and a real \(\delta \in (0, \vartheta/2)\), such that \[\rho\, I_k(F) \;<\; \Bigl(\frac{\vartheta}{2} - \delta\Bigr) \sum_{m=1}^{k} J_{k,\varepsilon}^{(m)}(F).\]
Uses: def_M_k, def_I_k, def_J_k, def_simplex
Proof. Write \(M := M_{k,\varepsilon}\), so that \(2\rho/\vartheta < M\) by hypothesis, and set \[\eta \;:=\; \frac{\vartheta}{2}\Bigl(M - \frac{2\rho}{\vartheta}\Bigr) \;>\; 0, \qquad\text{so that}\qquad \rho \;=\; \frac{\vartheta M}{2} - \eta .\] Apply Lemma 1 with tolerance \(\delta_0 := \min\bigl(\eta/\vartheta,\ M/2\bigr) > 0\): it produces a smooth \(F\) with closed support in \(\mathcal{T}_\varepsilon\), \(I_k(F) > 0\), and \[Q \;:=\; \frac{\sum_{m} J_{k,\varepsilon}^{(m)}(F)}{I_k(F)} \;>\; M - \delta_0 .\] Since \(\delta_0 \le M/2\) we get \(Q > M/2 > 0\).
Put \(\delta := \eta/(2Q)\), which is positive. From \(Q > M - \delta_0 \ge M - \eta/\vartheta\) we get \(\vartheta(M - Q) < \eta\), hence \[\frac{\vartheta M}{2} - \eta \;<\; \frac{\vartheta Q}{2} - \frac{\eta}{2} \;=\; \Bigl(\frac{\vartheta}{2} - \frac{\eta}{2Q}\Bigr) Q \;=\; \Bigl(\frac{\vartheta}{2} - \delta\Bigr) Q .\] Multiplying by \(I_k(F) > 0\) and using \(Q\, I_k(F) = \sum_m J_{k,\varepsilon}^{(m)}(F)\) gives \[\rho\, I_k(F) \;<\; \Bigl(\frac{\vartheta}{2} - \delta\Bigr) \sum_{m=1}^{k} J_{k,\varepsilon}^{(m)}(F).\] Finally \(\delta < \vartheta/2\): the displayed inequality has strictly positive left-hand side (as \(\rho > 0\) and \(I_k(F) > 0\)), and \(\sum_m J_{k,\varepsilon}^{(m)}(F) \ge 0\) since each term is an integral of a square, so the factor \(\vartheta/2 - \delta\) must be strictly positive. ◻
Proof uses: lem_Mk_approx, lem_density_smooth
Because \(M_{k,\varepsilon}\) is defined as a supremum, the hypothesis of Lemma 1 is equivalent to the existence of a single competitor \(F_0 \in L^2(\mathcal{T}_\varepsilon)\) with \(I_k(F_0) \ne 0\) and \(\sum_m J_{k,\varepsilon}^{(m)}(F_0) > (2\rho/\vartheta)\, I_k(F_0)\). Both cases use it: one supplies a polynomial competitor on \(\mathcal{R}_k\) at \(k = 105\), \(\rho = 2\), the other a certified competitor on \(\mathcal{T}_{1/25}\) at \(k = 50\), \(\rho = 1\).
Lower bounds for \(M_{k,\varepsilon}\): polynomial profiles
A lower bound \(M_{k,\varepsilon} > c\) is obtained by exhibiting a single competitor. On the standard simplex the competitors used are the profiles \(F_P\) of Definition 1 attached to symmetric polynomials \[P \;=\; \sum_{i=1}^{d} a_i\, (1 - P_1)^{b_i} P_2^{\,c_i}, \qquad (b_i, c_i) \in \mathcal{B}_{d'},\ a_i \in \mathbb{R},\] in the monomial family of Definition 1. For such profiles both \(I_k(F_P)\) and \(\sum_m J_k^{(m)}(F_P)\) are quadratic forms in the coefficient vector with explicit rational entries; the ratio is then a generalised Rayleigh quotient, whose supremum over the coefficients is the largest generalised eigenvalue. Verifying \(M_k > c\) therefore reduces to a finite rational computation.
Lemma 1 (Dirichlet’s integral over the simplex). For all \(k \in \mathbb{N}\) and all \(a, a_1, \dotsc, a_k \in \mathbb{Z}_{\ge 0}\), \[\int_{\mathcal{R}_k} \Bigl(1 - \sum_{i=1}^{k} t_i\Bigr)^{a} \prod_{i=1}^{k} t_i^{\,a_i} \, dt_1 \dotsm dt_k \;=\; \frac{a! \prod_{i=1}^{k} a_i!}{\bigl(k + a + \sum_{i=1}^{k} a_i\bigr)!} .\]
Uses: def_simplex
Proof. One proves the scaled form: for every real \(s \ge 0\), \[\int_{\mathcal{R}_k(s)} \Bigl(s - \sum_{i} t_i\Bigr)^{a} \prod_{i} t_i^{\,a_i} \, dt \;=\; s^{\,k + a + \sum_i a_i}\, \frac{a! \prod_i a_i!}{\bigl(k + a + \sum_i a_i\bigr)!},\] by induction on \(k\); the statement of the lemma is the case \(s = 1\). (The scaled form is also what is needed to integrate over \(\mathcal{T}_\varepsilon\) and \(\mathcal{S}_\varepsilon\).)
For \(k = 0\) both sides equal \(s^a\, a!/a! = s^a\). For the inductive step, the integrand is continuous on the compact set \(\mathcal{R}_{k+1}(s)\), hence integrable, and Fubini’s theorem lets us peel off the first coordinate: \[\int_{\mathcal{R}_{k+1}(s)} \!\!\cdots \;=\; \int_0^s t_1^{\,a_1} \biggl( \int_{\mathcal{R}_k(s - t_1)} \Bigl((s - t_1) - \sum_{i \ge 2} t_i\Bigr)^{a} \prod_{i \ge 2} t_i^{\,a_i} \, dt_2 \dotsm dt_{k+1} \biggr) dt_1 .\] By the inductive hypothesis applied with scale \(s - t_1\), the inner integral equals \((s - t_1)^{\,k + a + \sum_{i \ge 2} a_i}\, a! \prod_{i \ge 2} a_i! / (k + a + \sum_{i \ge 2} a_i)!\). The remaining one-dimensional integral is a Beta integral: for \(A, e \in \mathbb{Z}_{\ge 0}\) and \(s \ge 0\), \[\int_0^s (s - t)^{A}\, t^{e} \, dt \;=\; s^{\,A + e + 1}\, \frac{A!\, e!}{(A + e + 1)!},\] which follows from \(B(A+1, e+1) = A!\,e!/(A + e + 1)!\) after the substitution \(t = su\). Combining the two factorial expressions and simplifying gives the scaled formula in dimension \(k + 1\). ◻
Lemma 1 (Power-sum moments over the simplex). For all \(k \in \mathbb{N}\) and all \(a, b \in \mathbb{Z}_{\ge 0}\), \[\int_{\mathcal{R}_k} (1 - P_1)^{a}\, P_2^{\,b} \, dt_1 \dotsm dt_k \;=\; \frac{a!\; G_{b,2}(k)}{(k + a + 2b)!},\] where \(P_1\) and \(P_2\) are the power sums of Notation 1 and \(G_{b,j}\) is the polynomial of Definition 1.
Uses: def_simplex, not_power_sum, def_Gbj
Proof. Expand \(P_2^{\,b} = \bigl(\sum_i t_i^2\bigr)^b\) by the multinomial theorem: \[\Bigl(\sum_{i=1}^{k} t_i^{2}\Bigr)^{b} \;=\; \sum_{\substack{v \in \mathbb{Z}_{\ge 0}^{k} \\ v_1 + \dotsb + v_k = b}} \frac{b!}{\prod_i v_i!} \prod_{i=1}^{k} t_i^{\,2 v_i}.\] Integrating term by term over \(\mathcal{R}_k\) and applying Lemma 1 with exponents \(a_i = 2v_i\) (so \(\sum_i a_i = 2b\)) gives \[\int_{\mathcal{R}_k} (1 - P_1)^{a} P_2^{\,b} \;=\; \frac{a!}{(k + a + 2b)!} \sum_{\substack{v \in \mathbb{Z}_{\ge 0}^{k} \\ \sum_i v_i = b}} \frac{b!}{\prod_i v_i!} \prod_{i=1}^{k} (2 v_i)! .\] By Definition 1, \(G_{b,2}(k) = b! \sum_{\sum_i v_i = b} \prod_i (2 v_i)!/v_i!\), which is literally the displayed sum, since \(b!/\prod_i v_i!\) times \(\prod_i (2v_i)!\) equals \(b! \prod_i (2v_i)!/v_i!\). ◻
Proof uses: lem_monomial_integration
Lemma 1 (The \(L^2\) mass of a polynomial profile). Let \(k \in \mathbb{N}\), let \(d \in \mathbb{N}\), let \(a_1, \dotsc, a_d \in \mathbb{R}\) and let \((b_1, c_1), \dotsc, (b_d, c_d) \in \mathbb{Z}_{\ge 0}^2\), and let \(P\) satisfy \(P = \sum_{i=1}^{d} a_i (1 - P_1)^{b_i} P_2^{\,c_i}\). Let \(F_P\) be the profile of Definition 1. Then \[I_k(F_P) \;=\; \sum_{i=1}^{d} \sum_{j=1}^{d} a_i a_j\, \frac{(b_i + b_j)!\; G_{c_i + c_j,\,2}(k)}{\bigl(k + b_i + b_j + 2c_i + 2c_j\bigr)!} .\]
Uses: def_I_k, def_polynomial_F, def_basis_Bd, def_Gbj, not_power_sum
Proof. Since \(F_P\) vanishes off \(\mathcal{R}_k\) and agrees with \(P\) on it, \(I_k(F_P) = \int_{\mathcal{R}_k} P^2\). Expanding the square, \[P^2 \;=\; \sum_{i} \sum_{j} a_i a_j\, (1 - P_1)^{b_i + b_j}\, P_2^{\,c_i + c_j},\] using \((1 - P_1)^{b_i}(1 - P_1)^{b_j} = (1 - P_1)^{b_i + b_j}\) and likewise for \(P_2\). The integrand is continuous on the compact set \(\mathcal{R}_k\), so the finite sum may be integrated term by term, and each term is evaluated by Lemma 1 with \(a = b_i + b_j\) and \(b = c_i + c_j\). ◻
Proof uses: lem_integration_formula
Lemma 1 (The marginal mass of a polynomial profile). Let \(k,d \ge 1\), let \(a_1, \dotsc, a_d \in \mathbb{R}\), let \((b_1,c_1),\dotsc,(b_d,c_d) \in \mathbb{Z}_{\ge 0}^2\), let \(P = \sum_{i=1}^{d} a_i(1-P_1)^{b_i}P_2^{c_i}\), and let \(F_P\) be the profile of Definition 1. Then, for every \(m \in \{1, \dotsc, k\}\), \[J_k^{(m)}(F_P) \;=\; \sum_{i=1}^{d} \sum_{j=1}^{d} a_i a_j \sum_{c' = 0}^{c_i} \sum_{c'' = 0}^{c_j} \binom{c_i}{c'}\binom{c_j}{c''}\, \frac{\gamma(b_i, b_j, c_i, c_j, c', c'')\; G_{c' + c'',\,2}(k - 1)} {\bigl(k + b_i + b_j + 2c_i + 2c_j + 1\bigr)!},\] where \(\gamma\) is the function given by \[\gamma(b_i, b_j, c_i, c_j, c', c'') \;=\; \frac{b_i!\; b_j!\; (2c_i - 2c')!\; (2c_j - 2c'')!\; \bigl(b_i + b_j + 2c_i + 2c_j - 2c' - 2c'' + 2\bigr)!} {\bigl(b_i + 2c_i - 2c' + 1\bigr)!\; \bigl(b_j + 2c_j - 2c'' + 1\bigr)!} .\] In particular \(J_k^{(m)}(F_P)\) does not depend on \(m\).
Proof. By Lemma 1, \[J_k^{(m)}(F_P) \;=\; \int_{\mathcal{R}_k^{(m)}} \Bigl( \int_0^{\,1 - P_1'} P(\dotsc, s, \dotsc)\, ds \Bigr)^{\!2} \prod_{i \ne m} dt_i ,\] where \(P_1' = \sum_{i \ne m} t_i\) and \(s\) occupies the \(m\)-th slot. Since \(P\) is symmetric, the value of the inner integral is independent of which slot is singled out, and the outer domain \(\mathcal{R}_k^{(m)}\) is a copy of \(\mathcal{R}_{k-1}\); this gives the asserted independence of \(m\), and we may take \(m = 1\).
The inner integral. Writing \(P_2' = \sum_{i \ne m} t_i^2\), the \(i\)-th monomial of \(P\) evaluated at \((s, (t_i)_{i \ne m})\) is \(a_i (1 - P_1' - s)^{b_i}(s^2 + P_2')^{c_i}\). Binomially expanding \((s^2 + P_2')^{c_i} = \sum_{c' \le c_i} \binom{c_i}{c'} (P_2')^{c'} s^{2c_i - 2c'}\) and evaluating each one-dimensional integral by the Beta formula \(\int_0^L (L - s)^{B} s^{E}\, ds = L^{B + E + 1} B!\,E!/(B + E + 1)!\) with \(L = 1 - P_1'\) gives \[\int_0^{1 - P_1'} P \, ds \;=\; \sum_{i=1}^{d} a_i \sum_{c' = 0}^{c_i} \binom{c_i}{c'} (P_2')^{c'}\, (1 - P_1')^{\,b_i + 2c_i - 2c' + 1}\, \frac{b_i!\,(2c_i - 2c')!}{(b_i + 2c_i - 2c' + 1)!} .\]
Squaring. Multiplying this finite sum by itself produces the fourfold sum over \((i, j, c', c'')\), with the \((P_2')\)-powers combining to \((P_2')^{c' + c''}\), the \((1 - P_1')\)-powers combining to \((1 - P_1')^{\,b_i + b_j + 2c_i + 2c_j - 2c' - 2c'' + 2}\), and the factorial prefactors combining to \(\gamma(b_i, b_j, c_i, c_j, c', c'')\) divided by \(\bigl(b_i + b_j + 2c_i + 2c_j - 2c' - 2c'' + 2\bigr)!\).
The outer integral. Integrating term by term over \(\mathcal{R}_k^{(m)} \cong \mathcal{R}_{k-1}\) and applying Lemma 1 in dimension \(k - 1\) with \(a = b_i + b_j + 2c_i + 2c_j - 2c' - 2c'' + 2\) and \(b = c' + c''\) replaces each term by \[\frac{\bigl(b_i + b_j + 2c_i + 2c_j - 2c' - 2c'' + 2\bigr)!\; G_{c' + c'',\,2}(k-1)} {\bigl(k + b_i + b_j + 2c_i + 2c_j + 1\bigr)!},\] the denominator being \((k - 1) + a + 2b\) factorial. The factorial in the numerator cancels the one introduced in the squaring step, leaving exactly the stated expression. ◻
Proof uses: lem_J_iterated, lem_integration_formula
Lemma 1 (Positive definiteness of the Gram matrices). Let \((X, \mu)\) be a measure space, let \(f_1, \dotsc, f_d : X \to \mathbb{R}\) be measurable with \(f_i^2\) integrable, and suppose the family is linearly independent modulo null sets: whenever \(\alpha \in \mathbb{R}^d\) satisfies \(\sum_i \alpha_i f_i = 0\) \(\mu\)-almost everywhere, then \(\alpha = 0\). Let \(c > 0\), and let \(A\) be the symmetric matrix satisfying \[A \;=\; (A_{ij})_{1 \le i, j \le d}, \qquad A_{ij} \;=\; c \int_X f_i f_j \, d\mu .\] Then \(A\) is positive definite.
Proof. Symmetry is clear. For \(\alpha \in \mathbb{R}^d\), expanding the square and using integrability of each product \(f_i f_j\) (a consequence of Cauchy–Schwarz and the integrability of the \(f_i^2\)), \[\alpha^{\!\top} A \alpha \;=\; c \int_X \Bigl(\sum_{i=1}^{d} \alpha_i f_i\Bigr)^{\!2} d\mu \;\ge\; 0 .\] If \(\alpha^{\!\top} A \alpha = 0\) then, as \(c > 0\) and the integrand is non-negative, \(\sum_i \alpha_i f_i = 0\) almost everywhere, whence \(\alpha = 0\) by hypothesis. So the form is strictly positive on \(\mathbb{R}^d \setminus \{0\}\).
This is applied twice. Taking \(X = \mathcal{R}_k\) with Lebesgue measure and \(f_i\) the monomials \((1 - P_1)^{b_i} P_2^{\,c_i}\) shows that the matrix \(A_I\) whose quadratic form is \(I_k(F_P)\) (Lemma 1) is positive definite. Taking \(X = \mathcal{R}_k^{(m)}\) and \(f_i\) the marginals \(T_m f_i\) for one fixed \(m\) shows that the matrix with entries \(\int (T_m f_i)(T_m f_j)\) is positive definite; the remaining \(k - 1\) matrices in the sum \(A_J := \bigl(\sum_{m} \int (T_m f_i)(T_m f_j)\bigr)_{ij}\), whose quadratic form is \(\sum_m J_k^{(m)}(F_P)\) (Lemma 1), are positive semidefinite by the same computation, so \(A_J\) is positive definite. Only the independence of the marginals for a single \(m\) is needed. ◻
Proof uses: lem_Ik_quadratic, lem_quad_forms
Lemma 1 (The generalised Rayleigh quotient). Let \(d \ge 1\) and let \(A_1, A_2\) be real symmetric positive definite \(d \times d\) matrices. Then there exists \(\lambda \in \mathbb{R}\) with all of the following properties simultaneously:
\(\alpha^{\!\top} A_2 \alpha \le \lambda\, \alpha^{\!\top} A_1 \alpha\) for every \(\alpha \in \mathbb{R}^d\), with equality for some \(\alpha_0 \ne 0\);
\(\lambda = \sup\bigl\{ (\alpha^{\!\top} A_2 \alpha)/(\alpha^{\!\top} A_1 \alpha) : \alpha \ne 0 \bigr\}\);
\(A_2 \alpha = \lambda A_1 \alpha\) for some \(\alpha \ne 0\), and \(\nu \le \lambda\) whenever \(A_2 \alpha = \nu A_1 \alpha\) for some \(\alpha \ne 0\);
\(\lambda\) is the largest element of the spectrum of \(A_1^{-1} A_2\).
Proof. Since \(A_1\) is positive definite it factors as \(A_1 = L L^{\!\top}\) with \(L\) invertible. Put \(B := L^{-1} A_2 (L^{-1})^{\!\top}\); then \(B\) is symmetric, and it is positive definite because \(\beta^{\!\top} B \beta = \alpha^{\!\top} A_2 \alpha\) for \(\alpha := (L^{-1})^{\!\top}\beta\), and \(\beta \mapsto \alpha\) is a linear bijection. Substituting \(\beta := L^{\!\top} \alpha\) turns the generalised quotient into an ordinary one: \[\frac{\alpha^{\!\top} A_2 \alpha}{\alpha^{\!\top} A_1 \alpha} \;=\; \frac{\beta^{\!\top} B \beta}{\beta^{\!\top}\beta}, \qquad \alpha \ne 0 \iff \beta \ne 0 .\] By the spectral theorem \(B\) has an orthonormal eigenbasis with real eigenvalues; let \(\lambda\) be the largest of them and \(\beta_0\) a corresponding unit eigenvector. Expanding \(\beta\) in the eigenbasis gives \(\beta^{\!\top} B \beta \le \lambda\, \beta^{\!\top}\beta\) for all \(\beta\), with equality at \(\beta_0\); this is (i), and (ii) follows since the quotient is bounded above by \(\lambda\) and attains it.
For (iii), \(B\beta = \nu\beta\) with \(\beta \ne 0\) is equivalent, under \(\beta = L^{\!\top} \alpha\), to \(A_2 \alpha = \nu A_1 \alpha\) with \(\alpha \ne 0\); so the generalised eigenvalues of the pair \((A_1, A_2)\) are exactly the eigenvalues of \(B\), and \(\lambda\) is the largest. For (iv), \(A_2 \alpha = \nu A_1 \alpha\) is equivalent to \((A_1^{-1} A_2)\alpha = \nu \alpha\) because \(A_1\) is invertible, and for a square real matrix the spectrum consists precisely of its eigenvalues; so the spectrum of \(A_1^{-1} A_2\) is the set of generalised eigenvalues, whose largest element is \(\lambda\). ◻
The three preceding lemmas combine into the criterion by which every lower bound for \(M_k = M_{k,0}\) is certified.
Proposition 1 (Polynomial certificates for lower bounds on \(M_k\)). Let \(k \ge 1\) and \(d \ge 1\), let \((b_1, c_1), \dotsc, (b_d, c_d) \in \mathbb{Z}_{\ge 0}^2\) be pairwise distinct, and let \(\alpha = (a_1, \dotsc, a_d) \in \mathbb{R}^d\) be non-zero. Write \(A(\alpha)\) and \(B(\alpha)\) for the two quadratic forms of Lemmas 1 and 1, that is, \[A(\alpha) \;=\; \sum_{i,j} a_i a_j\, \frac{(b_i + b_j)!\; G_{c_i + c_j,\,2}(k)}{(k + b_i + b_j + 2c_i + 2c_j)!},\] \[B(\alpha) \;=\; \sum_{i,j} a_i a_j \!\!\sum_{c' = 0}^{c_i} \sum_{c'' = 0}^{c_j}\!\! \binom{c_i}{c'}\binom{c_j}{c''} \frac{\gamma(b_i, b_j, c_i, c_j, c', c'')\; G_{c' + c'',\,2}(k-1)} {(k + b_i + b_j + 2c_i + 2c_j + 1)!} ,\] with \(\gamma\) the explicit rational constant of Lemma 1. If \(c \in \mathbb{R}\) satisfies \(c\, A(\alpha) < k\, B(\alpha)\), then \(M_k > c\).
Uses: def_M_k, def_polynomial_F, def_basis_Bd, def_Gbj, not_power_sum, lem_Ik_quadratic, lem_quad_forms
Proof. Let \(P := \sum_i a_i (1 - P_1)^{b_i} P_2^{\,c_i}\) and let \(F_P\) be the associated profile, which vanishes outside \(\mathcal{R}_k = \mathcal{T}_0\). The monomials \((1 - P_1)^{b_i} P_2^{\,c_i}\), for pairwise distinct index pairs, are linearly independent as functions on \(\mathcal{R}_k\) modulo null sets; hence by Lemma 1 the Gram matrix \(A_I\) whose quadratic form is \(A(\alpha)\) is positive definite, and therefore \(A(\alpha) > 0\) because \(\alpha \ne 0\). By Lemma 1, \(I_k(F_P) = A(\alpha) > 0\), so \(F_P\) is an admissible competitor in Definition 1 at \(\varepsilon = 0\). By Lemma 1 each of the \(k\) marginal masses equals \(B(\alpha)\), so \(\sum_{m=1}^{k} J_k^{(m)}(F_P) = k\, B(\alpha)\). Definition 1 therefore gives \[M_k \;\ge\; \frac{k\, B(\alpha)}{A(\alpha)} \;>\; c ,\] the last step being the hypothesis divided by \(A(\alpha) > 0\).
The hypothesis is an inequality between two rational numbers whenever \(\alpha\) is rational, so it can be certified by an exact finite computation. The optimal constant obtainable this way from a fixed monomial family is, by Lemma 1 applied to the pair \((A_I, A_J)\) — both positive definite by Lemma 1 — equal to \(k\) times the largest element of the spectrum of \(A_I^{-1} A_J\). ◻
Proof uses: lem_gram_posdef, lem_rayleigh
The two explicit instances of this computation are carried out in the two later sections: on the standard simplex at \(k = 105\), with the \(42\) monomials indexed by \(\mathcal{B}_{11}\) and \(c = 4\), giving \(M_{105} > 4\); and on the \(\varepsilon\)-enlarged simplex at \(k = 50\), \(\varepsilon = 1/25\), with a larger basis, giving \(M_{50,\,1/25} > 4\). In each case the bound is used in Lemma 1.
Lemma 1 (The Beta integral). Let \(a, b \ge 1\) be integers. Then \[\int_0^1 t^{a-1} (1-t)^{b-1} \, dt \;=\; \frac{(a-1)!\,(b-1)!}{(a+b-1)!}.\]
Proof. The integral is the Euler Beta function \(B(a,b)\) evaluated at positive integer arguments. For integers it satisfies \(B(a,b) = \Gamma(a)\Gamma(b)/\Gamma(a+b)\), and \(\Gamma(n) = (n-1)!\) for \(n \ge 1\); substituting gives the stated quotient of factorials. ◻
Auxiliary arithmetic estimates
This section collects the arithmetic estimates used in the sieve analysis: the elementary totient and Möbius identities, the Mertens-type sums over integers coprime to a modulus, the evaluation of the truncated Mertens sum \(\sum_{u \le R,\,(u, W) = 1}\mu(u)^2/\varphi(u)\), the verification of the sieve-datum hypotheses for the two densities, the asymptotic for the summatory function \(H(z)\) of the sieve weight, and the partial-summation lemma that converts \(H\) into a \(G\)-weighted sum.
Throughout, \(\mathcal{W}\), \(V\), \(m\) and \(M\) denote generic moduli; \(W\) and \(R\) are the primorial and the truncation of Definitions 1 and 1, and \(\mathrm{e} = \exp(1)\) denotes the base of the natural logarithm (never a summation variable).
Elementary totient and Möbius identities
Lemma 1 (Multiplicativity formula for \(\varphi\) on a least common multiple). Let \(d\) and \(e\) be squarefree positive integers (each squarefree; the product \(de\) need not be). Then \[\varphi\bigl([d, e]\bigr)\,\varphi\bigl((d, e)\bigr) \;=\; \varphi(d)\,\varphi(e).\]
Uses: def_totient
Proof. Since \(d\) and \(e\) are squarefree, so are \([d, e]\) and \((d, e)\), and \[\varphi(n) \;=\; \prod_{p \mid n}(p - 1)\] for every squarefree \(n\). The prime divisors of \([d,e]\) are exactly those of \(d\) together with those of \(e\), while the prime divisors of \((d,e)\) are exactly those dividing both. Hence, counting each prime with the multiplicity with which it occurs on either side, \[\prod_{p \mid [d,e]}(p-1)\ \cdot \prod_{p \mid (d,e)}(p-1) \;=\; \prod_{p \mid d}(p-1)\ \cdot \prod_{p \mid e}(p-1),\] which is the assertion. ◻
Proof uses: def_totient
Lemma 1 (Reciprocal LCM splitting for \(\varphi\)). Let \(d\) and \(e\) be squarefree positive integers. Then \[\frac{1}{\varphi([d, e])} \;=\; \frac{1}{\varphi(d)\,\varphi(e)}\sum_{u \mid (d, e)} g(u),\] with \(g\) as in Definition 1.
Proof. By Lemma 1, \(\varphi([d,e]) = \varphi(d)\varphi(e)/\varphi((d,e))\), all the factors being positive. The integer \((d, e)\) is squarefree, so Lemma 1 applies to it and gives \(\varphi((d,e)) = \sum_{u \mid (d,e)} g(u)\). Substituting and inverting yields the claim. ◻
Proof uses: lem_phi_lcm_product, lem_phi_g_convolution
Lemma 1 (Reciprocal LCM splitting). For all positive integers \(d, e\) (no squarefreeness required), \[\frac{1}{[d, e]} \;=\; \frac{1}{de}\sum_{u \mid (d, e)} \varphi(u).\]
Uses: def_totient
Proof. The classical identity \(\sum_{u \mid n}\varphi(u) = n\), valid for every positive integer \(n\), applied at \(n = (d,e)\) turns the right-hand side into \((d,e)/(de)\). The claim is then the identity \([d,e]\,(d,e) = de\). ◻
Sublemma 1 (Divisor-sum form of \(1/\varphi\)). For every squarefree integer \(u \ge 1\), \[\frac{1}{\varphi(u)} \;=\; \frac{1}{u}\sum_{d \mid u}\frac{\mu(d)^2}{\varphi(d)}.\]
Uses: def_totient, def_mobius
Proof. For squarefree \(u\), \[\frac{u}{\varphi(u)} \;=\; \prod_{p \mid u}\frac{p}{p - 1} \;=\; \prod_{p \mid u}\Bigl(1 + \frac{1}{p - 1}\Bigr) \;=\; \sum_{d \mid u}\prod_{p \mid d}\frac{1}{p - 1} \;=\; \sum_{d \mid u}\frac{\mu(d)^2}{\varphi(d)}.\] The third equality expands the product over \(p \mid u\) as a sum over the subsets of the prime divisors of \(u\), i.e. over the divisors \(d \mid u\); the fourth uses \(\prod_{p \mid d}1/(p-1) = \mu(d)^2/\varphi(d)\) on squarefree \(d\). Dividing by \(u\) gives the claim. ◻
Sublemma 1 (Per-coordinate lcm-decomposition bijection). Let \(m\) be a squarefree positive integer, and let \(\mathcal{D}_m\) and \(\mathcal{C}_m\) be the sets \[\mathcal{D}_m = \{(d, e) : d, e \ge 1 \text{ squarefree},\ [d,e] = m\}, \qquad \mathcal{C}_m = \{(c, a, b) : c, a, b \ge 1 \text{ pairwise coprime},\ cab = m\}.\] Then \((d, e) \mapsto \bigl((d,e),\, d/(d,e),\, e/(d,e)\bigr)\) maps \(\mathcal{D}_m\) into \(\mathcal{C}_m\) and is a bijection, with inverse \((c, a, b) \mapsto (ca,\, cb)\).
Proof. Let \((d, e) \in \mathcal{D}_m\) and put \(c := (d,e)\), \(a := d/c\), \(b := e/c\). Then \(cab = (d,e)\,[d,e]/(d,e) \cdot \dotsb\); more directly, \(ca = d\), \(cb = e\) and \(cab = d\,b = d\,e/(d,e) = [d,e] = m\). The integers \(a\) and \(b\) are coprime because any common prime divisor would divide both \(d/c\) and \(e/c\) and hence, \(m\) being squarefree, already lie in \(c\); the same squarefreeness of \(m = cab\) forces \(c\) to be coprime to both \(a\) and \(b\). So the map lands in \(\mathcal{C}_m\).
Conversely let \((c, a, b) \in \mathcal{C}_m\) and put \(d := ca\), \(e := cb\). Both are divisors of the squarefree \(m\), hence squarefree, and \((d,e) = c\) and \([d,e] = cab = m\) by pairwise coprimality; so \((d,e) \in \mathcal{D}_m\). The two constructions are mutually inverse: starting from \((d,e)\) one recovers \(\bigl((d,e)\bigr)\cdot\bigl(d/(d,e)\bigr) = d\) and likewise for \(e\); starting from \((c,a,b)\) one recovers \((ca, cb) \mapsto (c, a, b)\) because \((ca, cb) = c\) when \((a,b) = 1\). ◻
Divisor sums, Mertens-type bounds and convergent series
Definition 1 (The von Mangoldt function \(\Lambda\)). Define \(\Lambda : \mathbb{N} \to \mathbb{R}_{\ge 0}\) by \[\Lambda(n) \;:=\; \begin{cases}\log p, & n = p^a \text{ for some prime } p \text{ and some } a \ge 1,\\ 0, & \text{otherwise.}\end{cases}\]
Lemma 1 (Mertens’ first theorem, von Mangoldt form). For every integer \(N \ge 1\), \[\Bigl|\sum_{d \le N}\frac{\Lambda(d)}{d} \;-\; \log N\Bigr| \;\le\; \log 4 + 5 .\]
Uses: def_von_mangoldt
Proof. Step 1: a Chebyshev upper bound. For every \(N \ge 0\), \[\sum_{d \le N}\Lambda(d) \;\le\; (\log 4 + 4)\,N .\] This is the classical elementary bound obtained from the central binomial coefficient: \(\prod_{N < p \le 2N}p\) divides \(\binom{2N}{N} \le 4^N\), so \(\sum_{N < p \le 2N}\log p \le N\log 4\), and summing the resulting dyadic estimate together with the \(O(\sqrt N \log N)\) contribution of the proper prime powers gives the stated linear bound with the explicit constant \(\log 4 + 4\).
Step 2: the factorial identity. Since every \(n \ge 1\) satisfies \(\log n = \sum_{d \mid n}\Lambda(d)\), grouping by \(d\) gives \[\sum_{n \le N}\log n \;=\; \sum_{d \le N}\Lambda(d)\Bigl\lfloor\frac{N}{d}\Bigr\rfloor .\] Replacing \(\lfloor N/d\rfloor\) by \(N/d\) costs at most \(\sum_{d \le N}\Lambda(d)\), which by Step 1 is at most \((\log 4 + 4)N\). Hence \[\Bigl|\sum_{n \le N}\log n \;-\; N\sum_{d \le N}\frac{\Lambda(d)}{d}\Bigr| \;\le\; (\log 4 + 4)\,N .\]
Step 3: Stirling-free evaluation of \(\sum_{n \le N}\log n\). Comparing the sum with \(\int_1^N \log t\,dt = N\log N - N + 1\) by monotonicity of \(\log\) gives \(\bigl|\sum_{n \le N}\log n - N\log N\bigr| \le N\). Combining with Step 2 and dividing by \(N\) yields \[\Bigl|\sum_{d \le N}\frac{\Lambda(d)}{d} - \log N\Bigr| \;\le\; \log 4 + 5 . \qedhere\] ◻
Proof uses: def_von_mangoldt
Lemma 1 (Mertens-type estimate for \(\sum \mu(u)^2\tau_k(u)/\varphi(u)\)). For every integer \(k \ge 1\) there is a constant \(C = C(k) > 0\) such that, for every real \(z \ge 2\), \[\sum_{u \le z} \frac{\mu(u)^2\, \tau_k(u)}{\varphi(u)} \;\le\; C\,(\log z)^k .\]
Uses: def_mobius, def_totient, def_tau_r
Proof. Rankin majorisation. For any non-negative \(a : \mathbb{N} \to \mathbb{R}_{\ge 0}\), any real \(Z \ge 1\) and any \(\kappa \ge 0\) one has \((Z/u)^\kappa \ge 1\) for \(u \le Z\), so \(\sum_{u \le Z}a(u) \le Z^\kappa \sum_{u \ge 1}a(u)u^{-\kappa}\). Apply this with \(a(u) := \mu(u)^2\tau_k(u)/\varphi(u)\) and \(\kappa := 1/\log(\mathrm{e}z) \in (0,1)\).
Euler factorisation. The majorant \(u \mapsto \mu(u)^2\tau_k(u)/(\varphi(u)u^\kappa)\) is non-negative, multiplicative and supported on squarefree integers; since \(\tau_k(p) = k\) and \(\varphi(p) = p-1\), and all prime-power terms with exponent \(\ge 2\) vanish, \[\sum_{u \ge 1}\frac{\mu(u)^2\tau_k(u)}{\varphi(u)u^\kappa} \;=\; \prod_p\Bigl(1 + \frac{k}{(p-1)p^\kappa}\Bigr).\]
Logarithmic bound. Using \(\log(1+x) \le x\) and splitting \(1/(p-1) = 1/p + 1/(p(p-1))\), \[\log\prod_p\Bigl(1 + \frac{k}{(p-1)p^\kappa}\Bigr) \;\le\; k\sum_p \frac{1}{p^{1+\kappa}} \;+\; k\sum_p\frac{1}{p(p-1)},\] and the second sum converges absolutely. For the first, start from Lemma 1: discarding the proper prime powers (whose total contribution to \(\sum_{d \le t}\Lambda(d)/d\) is \(O(1)\), being bounded by \(\sum_p \sum_{a \ge 2}(\log p)/p^a \ll \sum_p (\log p)/p^2\)) gives the one-sided prime form \(\sum_{p \le t}(\log p)/p \le \log t + O(1)\) for \(t \ge 2\). Abel summation of this against the weight \(1/\log t\) yields the elementary one-sided Mertens second bound \(\sum_{p \le t}1/p \le \log\log t + B\) with \(B\) absolute. A second Abel summation, now against the weight \(t^{-\kappa}\), gives \(\sum_p p^{-1-\kappa} = \log(1/\kappa) + O(1)\). With \(\kappa = 1/\log(\mathrm{e}z)\) this is \(\log\log(\mathrm{e}z) + O(1) \le \log\log z + O(1)\) for \(z \ge 3\) (the range \(2 \le z < 3\) being trivial). Only upper bounds on prime sums are used, so no Tauberian or Dirichlet-series input enters.
Conclusion. Exponentiating gives \(\prod_p(1 + k(p-1)^{-1}p^{-\kappa}) \le C_1(k)(\log z)^k\), while \(Z^\kappa = z^{1/\log(\mathrm{e}z)} \le \mathrm{e}\). Combining with the Rankin majorisation proves the lemma. ◻
Proof uses: lem_mertens_first
Lemma 1 (Truncated coprime divisor sum). For every integer \(k \ge 1\) there is a constant \(C = C(k) > 0\) such that, for every positive integer \(D\) and every real \(Z \ge 2\), \[\sum_{\substack{r \le Z\\ (r, D) = 1}} \frac{\mu(r)^2\,\tau_k(r)}{\varphi(r)} \;\le\; C\,(\log Z)^k .\] The constant is independent of \(D\).
Uses: def_mobius, def_totient, def_tau_r, lem_mertens_tau_k
Proof. The summand is non-negative, so discarding the coprimality condition only increases the sum; the claim is then Lemma 1. (The point of the statement is the uniformity in \(D\): the same constant serves for every modulus, which is what the downstream applications, where \(D\) varies with the sieve variables, require.) ◻
Proof uses: lem_mertens_tau_k
Lemma 1 (\(k\)-fold restricted reciprocal-totient sum). For every integer \(k \ge 1\) there is a constant \(C = C(k) > 0\) with the following property. For every real \(R \ge 2\), every modulus \(W \ge 1\) and every \(\mathbf{d} = (d_1, \dotsc, d_k) \in \mathbb{Z}_{\ge 1}^k\), \[\Bigl(\prod_{i=1}^k d_i\Bigr) \sum_{\substack{r_1, \dotsc, r_k \ge 1\\ d_i \mid r_i\ \forall i\\ \prod_i r_i \le R,\ \mu(\prod_i r_i)^2 = 1,\ (\prod_i r_i,\, W) = 1}} \ \prod_{i=1}^k \frac{1}{\varphi(r_i)} \;\le\; C\,(\log R)^k .\]
Uses: def_mobius, def_totient, lem_mertens_tau_k
Proof. Write \(r_i = d_i s_i\) and \(s' := \prod_i s_i\). Because \(\prod_i r_i\) is squarefree, the \(d_i\) are pairwise coprime and coprime to \(s'\), and \(\prod_i \varphi(r_i) = \varphi(\prod_i d_i) \varphi(s')\) by multiplicativity of \(\varphi\) on coprime factors. Hence the left-hand side is at most \[\frac{\prod_i d_i}{\varphi(\prod_i d_i)} \sum_{\substack{s' \le R/\prod_i d_i\\ \mu(s')^2 = 1,\ (s', W\prod_i d_i) = 1}} \frac{\tau_k(s')}{\varphi(s')},\] since a squarefree \(s'\) arises from at most \(\tau_k(s')\) tuples \((s_1, \dotsc, s_k)\) with \(\prod_i s_i = s'\). For the prefactor, the Dirichlet expansion of Lemma 1 gives \(\prod_i d_i/\varphi(\prod_i d_i) = \sum_{D \mid \prod_i d_i}\mu(D)^2/\varphi(D)\), which is bounded by \(\sum_{D \mid \prod_i d_i}\mu(D)^2\tau_k(D)/\varphi(D)\); and every prime dividing \(\prod_i d_i\) is at most \(R\). The remaining \(s'\)-sum is a truncated coprime divisor sum at modulus \(\bigl(W\prod_i d_i\bigr)\), so Lemma 1 bounds it by \(C(k)(\log R)^k\) with a constant that does not depend on that modulus — the uniformity there is exactly what is needed, since the modulus varies with \(\mathbf{d}\). Both factors are thus dominated by the truncated sum \(\sum_{u \le R}\mu(u)^2\tau_k(u)/\varphi(u)\) of Lemma 1 — more precisely, the whole expression is bounded by the Euler product \(\prod_{p \le R}\bigl(1 + k/(p-1)\bigr)\), and the argument of Lemma 1 bounds this by \(C(k)(\log R)^k\) for \(R \ge 2\). ◻
Lemma 1 (Convergent sum bound for \(\varphi\)). \(\displaystyle\sum_{s \ge 1}\frac{\mu(s)^2}{\varphi(s)^2} \;<\; \infty .\)
Uses: def_mobius, def_totient
Proof. The summand is non-negative, so it suffices to majorise it by a convergent series. The inequality \(s \le \varphi(s)\tau(s)\) established in the proof of Lemma 1 gives \(1/\varphi(s) \le \tau(s)/s\), whence \[\frac{\mu(s)^2}{\varphi(s)^2} \;\le\; \frac{\tau(s)^2}{s^2} \;=\; \frac{\tau(s)}{\sqrt s}\cdot\frac{\tau(s)}{s^{3/2}} \;\le\; \frac{2\,\tau(s)}{s^{3/2}},\] the last step by the elementary bound \(\tau(s) \le 2\sqrt s\) (divisors pair up as \(u \leftrightarrow s/u\) about \(\sqrt s\)). The Dirichlet series \(\sum_{s \ge 1}\tau(s)s^{-3/2} = \bigl(\sum_{s \ge 1}s^{-3/2}\bigr)^2\) converges since \(3/2 > 1\), so the partial sums of the original series are uniformly bounded, and a series of non-negative terms with bounded partial sums converges. ◻
Proof uses: lem_phi_lower_sf
Lemma 1 (Convergent sum bound for \(g\)). With \(g\) the multiplicative function of Definition 1 (\(g(p) = p - 2\)), \[\sum_{\substack{s \ge 1\\ (s, 2) = 1}} \frac{\mu(s)^2}{g(s)^2} \;<\; \infty,\] the summand being read as \(0\) whenever \(\mu(s)^2 = 0\); the sum is thus supported on odd squarefree \(s\), for which \(g(s) = \prod_{p \mid s}(p-2) > 0\).
Uses: def_mobius, def_g_func
Proof. The summand is non-negative and multiplicative, supported on odd squarefree integers. Its local factor at an odd prime \(p\) is \(1 + (p-2)^{-2}\), and at \(p = 2\) it is \(1\). Since \((p-2)^{-2} \le 9 p^{-2}\) for every odd prime \(p\) (indeed \(p - 2 \ge p/3\) for \(p \ge 3\)), the series \(\sum_{p}(p-2)^{-2}\) converges, so \(\prod_{p \text{ odd}}\bigl(1 + (p-2)^{-2}\bigr) < \infty\). Every partial sum over a finite set of integers is bounded by this Euler product, whence convergence. ◻
The constants \(A(M)\) and \(B(M)\)
Sublemma 1 (Euler product for \(\zeta(2)\)). With \(\zeta(2) = \sum_{n \ge 1}n^{-2}\), \[\zeta(2) \;=\; \prod_p\Bigl(1 - \frac{1}{p^2}\Bigr)^{\!-1},\] both the series and the product being convergent.
Proof. By unique factorisation, every \(n \ge 1\) has a unique representation \(n = \prod_p p^{a_p}\) with \(a_p \ge 0\) and only finitely many \(a_p\) non-zero, so \(n^{-2} = \prod_p p^{-2a_p}\). The series \(\sum_n n^{-2}\) converges absolutely (it is bounded by \(1 + \int_1^\infty t^{-2}\,dt = 2\)), which licenses reorganising it as a product over primes of the geometric series \(\sum_{a \ge 0}p^{-2a} = (1 - p^{-2})^{-1}\); the rearrangement is justified by letting \(P \to \infty\) in the partial product over \(p \le P\), whose expansion exhausts all \(n \ge 1\) in the limit. ◻
Definition 1 (The auxiliary Dirichlet sum \(A(M)\)). For every positive integer \(M\), set \[A(M) \;:=\; \sum_{\substack{e \ge 1\\ (e, M) = 1}}\frac{\mu(e)}{e^2}.\]
Uses: def_mobius
Sublemma 1 (Absolute convergence and the bound \(|A(M)| \le \zeta(2)\)). For every positive integer \(M\) the series of Definition 1 converges absolutely, and \(|A(M)| \le \zeta(2)\).
Uses: def_AM
Proof. Each summand satisfies \(|\mu(e)/e^2| \le 1/e^2\), so \[\sum_{\substack{e \ge 1\\ (e, M) = 1}}\Bigl|\frac{\mu(e)}{e^2}\Bigr| \;\le\; \sum_{e \ge 1}\frac{1}{e^2} \;=\; \zeta(2) \;<\; \infty .\] This is absolute convergence, and the triangle inequality applied to the (absolutely convergent) sum gives \(|A(M)| \le \zeta(2)\). ◻
Sublemma 1 (Closed form for \(A(M)\) in terms of \(\zeta(2)\)). For every positive integer \(M\) (squarefreeness is not required), \[A(M) \;=\; \frac{1/\zeta(2)}{\prod_{p \mid M}(1 - 1/p^2)} .\]
Proof. By Sublemma 1 the defining series converges absolutely, so it may be expanded as an Euler product. Its local factor at a prime \(p \mid M\) is \(1\) (only \(e = 1\) contributes at that prime), and at a prime \(p \nmid M\) it is \(\sum_{a \ge 0}\mu(p^a)p^{-2a} = 1 - p^{-2}\), since \(\mu(p^a) = 0\) for \(a \ge 2\). Hence \[A(M) \;=\; \prod_{p \nmid M}\Bigl(1 - \frac{1}{p^2}\Bigr).\] Splitting the product over all primes as \(\prod_p = \prod_{p \mid M}\cdot\prod_{p \nmid M}\) and invoking \(\prod_p(1 - p^{-2}) = 1/\zeta(2)\) (Sublemma 1) gives the stated closed form. Note that only the set of prime divisors of \(M\) enters, which is why no squarefreeness hypothesis is needed. ◻
Proof uses: slem_zeta_2_euler_product, slem_AM_bound_by_zeta2
Definition 1 (The auxiliary Dirichlet sum \(B(M)\)). For every positive integer \(M\), set \[B(M) \;:=\; \sum_{\substack{e \ge 1\\ (e, M) = 1}}\frac{\mu(e)\log e}{e^2}.\] The series converges absolutely, since \(|\mu(e)\log e / e^2| \le (\log e)/e^2\) and \(\sum_{e \ge 1}(\log e)/e^2 < \infty\).
Uses: def_mobius
Coprime harmonic and coprime-density sums
Definition 1 (The prime-cost quantity \(\ell_m\)). For a positive integer \(m\), set \[\ell_m \;:=\; \sum_{p \mid m}\frac{\log p}{p - 1},\] the sum running over the distinct prime divisors of \(m\).
Sublemma 1 (Coprime harmonic sum). There is an absolute constant \(C > 0\) such that, for every squarefree integer \(m \ge 1\) and every real \(Z \ge 1\), \[\Bigl|\sum_{\substack{n \le Z\\ (n, m) = 1}}\frac{1}{n} \;-\; \frac{\varphi(m)}{m}\bigl(\log Z + \gamma_0 + \ell_m\bigr)\Bigr| \;\le\; C\,\frac{\tau(m)}{Z},\] where \(\gamma_0\) is the Euler–Mascheroni constant.
Uses: def_ell_m, def_mobius, def_totient
Proof. Detect the coprimality condition by \(\mathbf{1}[(n,m) = 1] = \sum_{d \mid (n,m)}\mu(d)\) and swap the two finite sums: \[\sum_{\substack{n \le Z\\ (n,m)=1}}\frac1n \;=\; \sum_{d \mid m}\mu(d)\sum_{\substack{n \le Z\\ d \mid n}}\frac1n \;=\; \sum_{d \mid m}\frac{\mu(d)}{d}\sum_{n' \le Z/d}\frac{1}{n'} .\] The classical harmonic estimate \(\sum_{n' \le Y}1/n' = \log Y + \gamma_0 + O(1/Y)\) gives \[\sum_{\substack{n \le Z\\ (n,m)=1}}\frac1n \;=\; (\log Z + \gamma_0)\sum_{d \mid m}\frac{\mu(d)}{d} \;-\; \sum_{d \mid m}\frac{\mu(d)\log d}{d} \;+\; O\Bigl(\frac{\tau(m)}{Z}\Bigr),\] the error using \(\sum_{d \mid m}|\mu(d)| = \tau(m)\) for squarefree \(m\). Now \(\sum_{d \mid m}\mu(d)/d = \prod_{p \mid m}(1 - 1/p) = \varphi(m)/m\), and differentiating the identity \(\sum_{d \mid m}\mu(d)d^{-s} = \prod_{p \mid m}(1 - p^{-s})\) at \(s = 1\) gives \[-\sum_{d \mid m}\frac{\mu(d)\log d}{d} \;=\; \prod_{p \mid m}\Bigl(1 - \frac1p\Bigr)\sum_{p \mid m}\frac{p^{-1}\log p}{1 - p^{-1}} \;=\; \frac{\varphi(m)}{m}\,\ell_m .\] Combining the three displays proves the claim, with an absolute implied constant. ◻
Lemma 1 (Mertens sum \(\sum 1/n\) restricted to \((n, \mathcal{W}) = 1\)). There is an absolute constant \(C > 0\) with the following property. Let \(\mathcal{W} \ge 1\) be a squarefree integer and let \(x\) be a real number with \[x \;\ge\; \frac{\mathcal{W}\,\tau(\mathcal{W})}{\varphi(\mathcal{W})} .\] Then \[\Bigl|\sum_{\substack{n \le x\\ (n, \mathcal{W}) = 1}}\frac{1}{n} \;-\; \frac{\varphi(\mathcal{W})}{\mathcal{W}}\log x\Bigr| \;\le\; C\,\frac{\varphi(\mathcal{W})}{\mathcal{W}}\bigl(1 + \log\log(\mathrm{e}\mathcal{W})\bigr).\]
Proof. Sublemma 1 applied with \(m = \mathcal{W}\) and \(Z = x\) gives \[\sum_{\substack{n \le x\\ (n, \mathcal{W}) = 1}}\frac1n \;=\; \frac{\varphi(\mathcal{W})}{\mathcal{W}}\bigl(\log x + \gamma_0 + \ell_{\mathcal{W}}\bigr) \;+\; O\Bigl(\frac{\tau(\mathcal{W})}{x}\Bigr).\] The hypothesis on \(x\) is exactly \(\tau(\mathcal{W})/x \le \varphi(\mathcal{W})/\mathcal{W}\), so the error term is \(O(\varphi(\mathcal{W})/\mathcal{W})\); the same bound covers the constant \(\gamma_0\,\varphi(\mathcal{W})/\mathcal{W}\). It remains to bound \(\ell_{\mathcal{W}} = \sum_{p \mid \mathcal{W}}(\log p)/(p-1)\). Split the prime divisors of \(\mathcal{W}\) at \(1 + \log \mathcal{W}\). For \(p > 1 + \log\mathcal{W}\), use \(\sum_{p \mid \mathcal{W}}\log p = \log \mathcal{W}\) (valid since \(\mathcal{W}\) is squarefree) and \(p - 1 > \log\mathcal{W}\) to bound that part by \(1\); more crudely, the contribution is \(O(1)\). For \(p \le 1 + \log\mathcal{W}\), use \((\log p)/(p-1) \le 2(\log p)/p\) and the elementary Mertens bound \(\sum_{p \le t}(\log p)/p \le \log t + O(1)\) of Lemma 1 with \(t = 1 + \log\mathcal{W}\), which gives \(O\bigl(\log(1 + \log \mathcal{W})\bigr) = O\bigl(1 + \log\log(\mathrm{e}\mathcal{W})\bigr)\). Hence \(\ell_{\mathcal{W}} \ll 1 + \log\log(\mathrm{e}\mathcal{W})\), and collecting all error contributions proves the lemma. The normalisation \(1 + \log\log(\mathrm{e}\mathcal{W})\) rather than \(\log\log\mathcal{W}\) is used so that the error term is defined and non-negative for every \(\mathcal{W} \ge 1\); for large \(\mathcal{W}\) the two agree up to an additive constant. ◻
Proof uses: slem_coprime_harmonic, def_ell_m, lem_mertens_first
Definition 1 (The coprime density partial sum \(A_m(x)\)). For a positive integer \(m\) and a real number \(x\), set \[A_m(x) \;:=\; \sum_{\substack{f \le x\\ (f, m) = 1}}\frac{\mu(f)^2}{\varphi(f)} ,\] the sum running over the positive integers \(f \le x\) coprime to \(m\); thus \(A_m(x) = 0\) for \(x < 1\).
Uses: def_mobius, def_totient
Sublemma 1 (Coprime density sum). There is an absolute constant \(C > 0\) such that, for every squarefree integer \(m \ge 1\) and every real \(x \ge 2\), \[\Bigl|A_m(x) \;-\; \frac{\varphi(m)}{m}\bigl(\log x + \ell_m\bigr)\Bigr| \;\le\; C\Bigl(\frac{\varphi(m)}{m} \;+\; \frac{\tau(m)\log(2mx)}{\sqrt{x}}\Bigr).\]
Uses: def_A_m_x, def_ell_m, slem_coprime_harmonic
Proof. Step 1: a Dirichlet cofactor. Let \(\kappa\) be the multiplicative function defined by the Dirichlet convolution identity \(\mu^2/\varphi= \kappa \ast (1/\mathrm{id})\), i.e. \[\frac{\mu(f)^2}{\varphi(f)} \;=\; \sum_{de = f}\kappa(d)\,\frac{1}{e} .\] Comparing local Euler factors shows \(\kappa(1) = 1\), \(\kappa(p) = 1/(p(p-1))\), \(\kappa(p^2) = -1/(p(p-1))\) and \(\kappa(p^j) = 0\) for \(j \ge 3\); in particular \(\kappa\) is supported on cube-free integers, its local factor \(1 + \kappa(p) + \kappa(p^2)\) equals \(1\) at every prime, and \[\sum_{d \ge 1}|\kappa(d)|\,\sqrt{d} \;<\; \infty, \qquad \sum_{d \ge 1}|\kappa(d)|\log(2d) \;<\; \infty .\] Both convergence statements follow from the Euler-product bound \(\sum_{j \ge 0}|\kappa(p^j)|p^{j/2} \le \exp\bigl(2p^{-3/2}\bigr)\), together with the elementary partial-summation bound \(\sum_{d \le X}|\kappa(d)|\,d \ll \sqrt X \log(2X)\), which in turn comes from writing a cube-free \(d\) uniquely as \(b^2 a\) with \(a\) squarefree, \((a, b) = 1\), and using \(\sum_{b \le \sqrt X}1/\varphi(b) \ll 1 + \log X\). Consequently, for every \(Y \ge 1\), \[\sum_{d \ge Y}|\kappa(d)| \;\ll\; \frac{\log(2Y)}{\sqrt Y}, \qquad \sum_{d \ge Y}|\kappa(d)|\log(d/Y) \;\ll\; \frac{\log(2Y)}{\sqrt Y}\] (a dyadic decomposition of the range \(d \ge Y\)).
Step 2: the hyperbola identity. Summing \(\mu^2/\varphi= \kappa \ast (1/\mathrm{id})\) over \(f \le x\) with \((f, m) = 1\) and noting that \((f, m) = 1\) is equivalent to \((d, m) = (e, m) = 1\) for \(f = de\), \[A_m(x) \;=\; \sum_{\substack{d \le x\\ (d, m) = 1}}\kappa(d) \sum_{\substack{e \le x/d\\ (e, m) = 1}}\frac1e .\]
Step 3: insert the harmonic asymptotic. Apply Sublemma 1 to the inner sum for each \(d\), obtaining the main term \(\frac{\varphi(m)}{m}\bigl(\log(x/d) + \gamma_0 + \ell_m\bigr)\) and an error \(O(\tau(m)d/x)\); the errors sum, against \(\sum_{d \le x}|\kappa(d)|d \ll \sqrt x\log(2x)\), to \(O\bigl(\tau(m)\log(2x)/\sqrt x\bigr)\). For the main term, complete the \(d\)-sum to all \(d \ge 1\) coprime to \(m\); by Step 1 the discarded tail \(d > x\) contributes \(O\bigl(\tau(m)\log(2mx)/\sqrt x\bigr)\) (using \(\ell_m \le \log m\) to absorb the \(m\) into the logarithm). Since \(\kappa\) has local mass \(1\) at every prime, \(\sum_{(d,m)=1}\kappa(d) = 1\), and the completed main term is \[\frac{\varphi(m)}{m}\bigl(\log x + \gamma_0 + \ell_m\bigr) \;-\; \frac{\varphi(m)}{m}\sum_{\substack{d \ge 1\\(d,m)=1}}\kappa(d)\log d .\] The last sum is \(O(1)\) by the second convergence statement of Step 1, and \(\gamma_0\varphi(m)/m = O(\varphi(m)/m)\). Collecting all contributions gives the stated bound, with an absolute constant. ◻
Proof uses: slem_coprime_harmonic, def_ell_m, def_A_m_x
The sum \(T_d\) and the Mertens sum over integers coprime to \(W\)
Definition 1 (The auxiliary sum \(T_d\)). Let \(R \ge 1\) be real and let \(W \ge 1\) and \(d \ge 1\) be integers. Set \[T_d \;:=\; \sum_{\substack{v \le R/d\\ (v, dW) = 1}}\frac{\mu(v)^2}{v} .\] The sum is empty, and hence \(T_d = 0\), once \(d > R\).
Uses: def_mobius
Sublemma 1 (Outer-sum decomposition for \(\sum \mu^2/\varphi\)). For every integer \(R \ge 1\) and every integer \(W \ge 1\), \[\sum_{\substack{u \le R\\ (u, W) = 1}}\frac{\mu(u)^2}{\varphi(u)} \;=\; \sum_{\substack{d \le R\\ (d, W) = 1}}\frac{\mu(d)^2}{d\,\varphi(d)}\,T_d ,\] with \(T_d\) as in Definition 1.
Uses: def_T_d, def_mobius, def_totient
Proof. Substitute the identity \(1/\varphi(u) = (1/u)\sum_{d \mid u}\mu(d)^2/\varphi(d)\) of Sublemma 1 — legitimate because \(\mu(u)^2 \ne 0\) restricts the outer sum to squarefree \(u\) — and swap the two finite summations: \[\sum_{\substack{u \le R\\ (u, W) = 1}}\frac{\mu(u)^2}{\varphi(u)} \;=\; \sum_{\substack{d \le R\\ (d, W) = 1}}\frac{\mu(d)^2}{\varphi(d)} \sum_{\substack{u \le R,\ d \mid u\\ (u, W) = 1}}\frac{\mu(u)^2}{u} .\] Fix a squarefree \(d\) with \((d, W) = 1\) and write \(u = dv\) with \(v \le R/d\). On the support of \(\mu(u)^2\) the integer \(u\) is squarefree, which for \(u = dv\) is equivalent to \(v\) squarefree with \((v, d) = 1\); and \((u, W) = 1\) together with \((d, W) = 1\) is equivalent to \((v, W) = 1\). So the conditions on \(u\) become “\(v\) squarefree and \((v, dW) = 1\)”. Since \(\mu(u)^2/u = \mu(v)^2/(dv)\) on that support, the inner sum equals \(T_d/d\), which is the claim. ◻
Proof uses: slem_mu_phi_divisor_sum
Sublemma 1 (Möbius squarefree-inversion step for \(T_d\)). For every real \(R \ge 1\), every integer \(W \ge 1\) and every integer \(d \ge 1\) (no squarefreeness or coprimality is required), \[T_d \;=\; \sum_{\substack{e \le \sqrt{R/d}\\ (e, dW) = 1}}\frac{\mu(e)}{e^2} \sum_{\substack{w \le R/(de^2)\\ (w, dW) = 1}}\frac{1}{w} .\]
Uses: def_T_d, def_mobius
Proof. Start from the squarefree-detection identity \(\mu(v)^2 = \sum_{e^2 \mid v}\mu(e)\), valid for every positive integer \(v\): both sides are multiplicative, and at a prime power \(p^a\) the right-hand side is \(1\) for \(a \in \{0, 1\}\) and \(1 + \mu(p) = 0\) for \(a \ge 2\), matching \(\mu(p^a)^2\). Insert it into \(T_d = \sum_{v \le R/d,\,(v, dW) = 1}\mu(v)^2/v\) and swap the two finite sums. Writing \(v = e^2 w\) turns the constraint \(e^2 \mid v\) into a free \(w\)-sum with \(w \le R/(de^2)\), and the constraint \((v, dW) = 1\) into \((e, dW) = (w, dW) = 1\); the \(e\)-range is \(e \le \sqrt{R/d}\) because \(e^2 \le v \le R/d\). Since \(1/v = (1/e^2)(1/w)\), the displayed identity follows. (When \((e, dW) > 1\) the inner sum is empty, so restricting the \(e\)-sum to \((e, dW) = 1\) changes nothing.) ◻
Sublemma 1 (Inner \(w\)-sum via Mertens for \(e \le R^{1/8}\)). There is a constant \(C > 0\) such that, for all sufficiently large \(N\), the following holds for every squarefree \(d\) with \((d, W) = 1\) and \(d \le R^{1/2}\) and every integer \(e\) with \(1 \le e \le R^{1/8}\) and \((e, dW) = 1\). Let \(\mathcal{W}=dW\). Then \[\Bigl|\sum_{\substack{w \le R/(de^2)\\ (w, \mathcal{W}) = 1}}\frac{1}{w} \;-\; \frac{\varphi(\mathcal{W})}{\mathcal{W}}\bigl(\log R - \log d - 2\log e\bigr)\Bigr| \;\le\; C\,\frac{\varphi(\mathcal{W})}{\mathcal{W}} \bigl(1 + \log\log(\mathrm{e}\mathcal{W})\bigr).\]
Proof. Since \(d\) is squarefree, \(W\) is squarefree (Definition 1) and \((d, W) = 1\), the modulus \(\mathcal{W} = dW\) is a positive squarefree integer. We apply Lemma 1 with this modulus and with \(x := R/(de^2)\), and must verify \(x \ge \mathcal{W}\tau(\mathcal{W})/\varphi(\mathcal{W})\).
Lower bound on \(x\). From \(d \le R^{1/2}\) and \(e \le R^{1/8}\) we get \(de^2 \le R^{1/2}\cdot R^{1/4} = R^{3/4}\), hence \(x = R/(de^2) \ge R^{1/4}\).
Upper bound on \(\mathcal{W}\tau(\mathcal{W})/\varphi(\mathcal{W})\). For squarefree \(n\) the quantity \(n\tau(n)/\varphi(n)\) factors over the primes dividing \(n\), each contributing \(2p/(p-1) \le 4\); hence \(\mathcal{W}\tau(\mathcal{W})/\varphi(\mathcal{W}) \le 4^{\omega(\mathcal{W})}\). Next, for any fixed integer \(B \ge 2\), splitting the prime divisors of \(n\) at \(B\) gives \(\omega(n) \le \pi(B) + \log n/\log B\), hence \(4^{\omega(n)} \le 4^{\pi(B)}\,n^{\log 4/\log B}\). Choose \(B := 4^8 = 65536\), so that \(\log 4/\log B = 1/8\) and, with \(K := 4^{\pi(B)}\) an absolute constant, \[\frac{\mathcal{W}\tau(\mathcal{W})}{\varphi(\mathcal{W})} \;\le\; K\,\mathcal{W}^{1/8}.\] By Lemma 1 we have \(W \ll (\log\log N)^2\), and \(d \le R^{1/2}\), so \(\mathcal{W} = dW \le R^{1/2}(\log\log N)^2\) and therefore \[\frac{\mathcal{W}\tau(\mathcal{W})}{\varphi(\mathcal{W})} \;\le\; K\,R^{1/16}\,(\log\log N)^{1/4} .\] Since \(R = N^{\vartheta/2 - \delta}\) is a fixed positive power of \(N\), the right-hand side is \(\le R^{1/4} \le x\) for all sufficiently large \(N\), the factor \((\log\log N)^{1/4}\) being of lower order than \(R^{3/16}\). This verifies the hypothesis.
With the hypothesis verified, Lemma 1 gives \[\Bigl|\sum_{\substack{w \le R/(de^2)\\ (w, \mathcal{W}) = 1}}\frac1w - \frac{\varphi(\mathcal{W})}{\mathcal{W}}\log\frac{R}{de^2}\Bigr| \le C\,\frac{\varphi(\mathcal{W})}{\mathcal{W}}\bigl(1 + \log\log(\mathrm{e}\mathcal{W})\bigr),\] and \(\log\bigl(R/(de^2)\bigr) = \log R - \log d - 2\log e\). ◻
Proof uses: lem_mertens_reciprocal_W, lem_W_size
Sublemma 1 (Evaluation of \(T_d\)). There is a constant \(C > 0\) such that, for all sufficiently large \(N\) and every squarefree \(d\) with \((d, W) = 1\) and \(d \le R^{1/2}\). Let \(\mathcal{W}=dW\). Then \[\Bigl|T_d \;-\; \frac{\varphi(\mathcal{W})}{\mathcal{W}} \bigl[A(\mathcal{W})\log(R/d) \;-\; 2B(\mathcal{W})\bigr]\Bigr| \;\le\; C\,\frac{\varphi(\mathcal{W})}{\mathcal{W}}\bigl(\log(2d) + \log\log W\bigr) \;+\; C\,\frac{\log R}{R^{1/8}} .\]
Uses: def_T_d, def_R, def_W_trick, def_AM, def_BM
Proof. Sublemma 1 rewrites \(T_d\) as an \(e\)-sum weighted by \(\mu(e)/e^2\) times an inner \(w\)-sum. Split the \(e\)-range at \(e = R^{1/8}\).
Range \(e \le R^{1/8}\). Sublemma 1 evaluates the inner sum for each such \(e\), with error \(O\bigl(\frac{\varphi(\mathcal{W})}{\mathcal{W}}(1 + \log\log(\mathrm{e}\mathcal{W}))\bigr)\). Summing the main term against \(\mu(e)/e^2\) over \(e \le R^{1/8}\) coprime to \(\mathcal{W}\) produces \[\frac{\varphi(\mathcal{W})}{\mathcal{W}}\Bigl[\log(R/d)\sum_{\substack{e \le R^{1/8}\\(e,\mathcal{W})=1}}\frac{\mu(e)}{e^2} \;-\; 2\sum_{\substack{e \le R^{1/8}\\(e,\mathcal{W})=1}}\frac{\mu(e)\log e}{e^2}\Bigr].\] Completing both \(e\)-sums to infinity replaces them by \(A(\mathcal{W})\) and \(B(\mathcal{W})\) of Definitions 1 and 1; by the tail bounds \(\bigl|A(M) - \sum_{e \le m}\bigr| \le 1/m\) and \(\bigl|B(M) - \sum_{e \le m}\bigr| \le (\log m + 1)/m\) (each a comparison with \(\sum_{e > m}e^{-2}\) and \(\sum_{e > m}(\log e)e^{-2}\) respectively), applied at \(m = \lfloor R^{1/8}\rfloor\), the cost is \(O(\log R \cdot R^{-1/8})\). The accumulated error terms contribute \(O\bigl(\frac{\varphi(\mathcal{W})}{\mathcal{W}}(1 + \log\log(\mathrm{e}\mathcal{W}))\bigr)\) after summation against \(\sum_e e^{-2} \le 2\).
Range \(e > R^{1/8}\). Here use the trivial bounds \(|\mu(e)/e^2| \le 1/e^2\) and \(\sum_{w \le R/(de^2)}1/w \le 1 + \log R\), so this part is \[\ll \log R \sum_{e > R^{1/8}}\frac{1}{e^2} \;\ll\; \frac{\log R}{R^{1/8}} .\]
Normalising the error. Finally \(\log\log(\mathrm{e}\mathcal{W}) = \log\log(\mathrm{e}dW) \ll \log(2d) + \log\log W\) for \(W \ge \mathrm{e}\), which converts the middle error into the stated shape. Adding the three contributions gives the sublemma. ◻
Proof uses: slem_T_d_mobius, slem_T_d_inner_w_sum, def_AM, def_BM
Sublemma 1 (Tail bound for \(d > R^{1/2}\)). There is a constant \(C > 0\) such that, for all sufficiently large \(N\), \[\Bigl|\sum_{\substack{d > \sqrt R,\ d\text{ squarefree}\\ (d, W) = 1}} \frac{\mu(d)^2}{d\,\varphi(d)}\,T_d\Bigr| \;\le\; C\,\frac{(\log R)^3}{\sqrt R} .\]
Uses: def_T_d, def_R, def_W_trick
Proof. From \(\mu(v)^2/v \le 1/v\) and \(\sum_{v \le R/d}1/v \le 1 + \log R\) we get \(0 \le T_d \le 1 + \log R \ll \log R\) for \(R \ge 2\), uniformly in \(d\) and \(W\). By Lemma 1, \(1/(d\varphi(d)) \le \tau(d)/d^2\). Lemma 1 with \(k = 2\) gives \(\sum_{d \le X}\mu(d)^2\tau(d)/\varphi(d) \ll (\log X)^2\), hence also \(\sum_{d \le X}\mu(d)^2\tau(d)/d \ll (\log X)^2\); partial summation converts this into the tail estimate \[\sum_{d > Y}\frac{\mu(d)^2\tau(d)}{d^2} \;\ll\; \frac{(\log Y)^2 + 1}{Y}.\] Applying it with \(Y = \sqrt R\) bounds \(\sum_{d > \sqrt R}\mu(d)^2/(d\varphi(d)) \ll (\log R)^2/\sqrt R\), and multiplying by the uniform bound \(T_d \ll \log R\) gives the claim. ◻
Proof uses: lem_phi_lower_sf, lem_mertens_tau_k
Sublemma 1 (Leading-coefficient telescoping identity). For every integer \(W \ge 1\), \[\sum_{\substack{d \ge 1\\ (d, W) = 1}}\frac{\mu(d)^2}{d^2}\,A(dW) \;=\; 1 .\]
Proof. Write \(f(d) := \mu(d)^2 A(dW)/d^2\). By Sublemma 1, \(A(dW) = (1/\zeta(2))/\prod_{p \mid dW}(1 - p^{-2})\), so for squarefree \(d\) coprime to \(W\) \[f(d) \;=\; A(W)\prod_{p \mid d}\frac{1/p^2}{1 - 1/p^2} .\] In particular \(f(1) = A(W)\), and \(f\) is multiplicative up to the normalising constant \(A(W)\): for coprime squarefree \(d_1, d_2\) both coprime to \(W\) one has \(f(d_1 d_2)A(W) = f(d_1)f(d_2)\); and at a prime \(p \nmid W\), \(f(p) = p^{-2}A(W)/(1 - p^{-2})\). Since the values of \(f\) at distinct primes multiply in this normalised sense, and \(\mu(d)^2\) restricts the sum to squarefree \(d\), each prime \(p \nmid W\) contributes independently either \(1\) or \(p^{-2}(1-p^{-2})^{-1}\), giving \[\sum_{\substack{d \ge 1\\(d, W) = 1}}\frac{\mu(d)^2}{d^2}A(dW) \;=\; A(W)\prod_{p \nmid W}\Bigl(1 + \frac{1}{p^2}\cdot\frac{1}{1 - 1/p^2}\Bigr).\] (The rearrangement is licensed by absolute convergence: \(|A(dW)| \le \zeta(2)\) by Sublemma 1 and \(\sum_d d^{-2} < \infty\).) The elementary identity \(1 + x/(1-x) = 1/(1-x)\) with \(x = p^{-2}\) turns the local factor into \((1 - p^{-2})^{-1}\), so the product equals \(\prod_{p \nmid W}(1 - p^{-2})^{-1}\). Combining with \(A(W) = (1/\zeta(2))/\prod_{p \mid W}(1 - p^{-2})\), \[\sum_{\substack{d \ge 1\\(d, W) = 1}}\frac{\mu(d)^2}{d^2}A(dW) \;=\; \frac{1/\zeta(2)}{\prod_{p \mid W}(1 - p^{-2})\prod_{p \nmid W}(1 - p^{-2})} \;=\; \frac{1/\zeta(2)}{\prod_p (1 - p^{-2})} \;=\; 1,\] the last equality by Sublemma 1. ◻
Proof uses: slem_AM_value, slem_AM_bound_by_zeta2, slem_zeta_2_euler_product
Lemma 1 (Mertens sum over \(u\) coprime to \(W\)). Let \(R\) and \(W\) be as in Definitions 1 and 1. There is a constant \(C > 0\) such that, for all sufficiently large \(N\), \[\Bigl|\sum_{\substack{u \le R\\ (u, W) = 1}}\frac{\mu(u)^2}{\varphi(u)} \;-\; \frac{\varphi(W)}{W}\log R\Bigr| \;\le\; C\,\frac{\varphi(W)}{W}\,\log\log W .\]
Uses: def_R, def_W_trick, def_mobius, def_totient
Proof. By Sublemma 1 the left-hand sum equals \(\sum_d \mu(d)^2 T_d/(d\varphi(d))\). Sublemma 1 bounds the contribution of \(d > \sqrt R\) by \(O\bigl((\log R)^3/\sqrt R\bigr)\), which is \(o(\varphi(W)/W)\) for large \(N\) since \(\varphi(W)/W \gg 1/\log\log N\) by Lemma 1. Restrict to \(d \le \sqrt R\).
Insert the expansion of Sublemma 1 and use \(\varphi(dW)/(dW) = (\varphi(d)/d)(\varphi(W)/W)\), valid because \((d, W) = 1\). This gives \[\sum_{\substack{u \le R\\ (u, W) = 1}}\frac{\mu(u)^2}{\varphi(u)} \;=\; \frac{\varphi(W)}{W}\Bigl[(\log R)\,\mathcal{S}_1 - \mathcal{S}_2 - 2\mathcal{S}_3\Bigr] \;+\; O\Bigl(\frac{\varphi(W)}{W}\log\log W\Bigr),\] where, with all sums over squarefree \(d\) coprime to \(W\), \[\mathcal{S}_1 := \sum_{d}\frac{\mu(d)^2 A(dW)}{d^2},\qquad \mathcal{S}_2 := \sum_{d}\frac{\mu(d)^2 (\log d)A(dW)}{d^2},\qquad \mathcal{S}_3 := \sum_{d}\frac{\mu(d)^2 B(dW)}{d^2},\] and the error term also absorbs the \(\log(2d)\) contributions of Sublemma 1 (which sum, against \(\mu(d)^2\varphi(d)/(d^2\varphi(d)) \le d^{-2}\), to \(O(1)\)) and the \(O(\log R/R^{1/8})\) terms. By Sublemma 1, \(\mathcal{S}_1 = 1\) once the \(d\)-sum is completed, and completing it costs \(O(R^{-1/2})\) by Sublemma 1. The sums \(\mathcal{S}_2\) and \(\mathcal{S}_3\) are \(O(1)\) in absolute value: each is bounded by \(\zeta(2)\sum_{d \ge 1}(1 + \log d)/d^2 \ll 1\), again using \(|A(M)|, |B(M)| \ll 1\) from Sublemma 1 and the corresponding trivial bound for \(B\). Assembling and absorbing the \(O(1)\) and \(O(R^{-1/2}\log R)\) contributions into \(O\bigl(\frac{\varphi(W)}{W}\log\log W\bigr)\) — legitimate since \(\log\log W \ge 1\) for \(N\) large — gives the lemma. ◻
The two sieve densities
Throughout this subsection \((\gamma, V)\) denotes a sieve datum in the sense of Definition 1, with parameters \(A_1, A_2, A_3, L\), and \(g_*\) is the derived totally multiplicative function of Definition 1. We write \[\Delta(\gamma; w, z) \;:=\; \sum_{w \le p \le z}\frac{\gamma(p)\log p}{p} \;-\; \log(z/w) \qquad (2 \le w \le z)\] for the Mertens deviation of a density \(\gamma\).
Sublemma 1 (\(g_*\) for the Maynard datum). For the Maynard sieve datum (Definition 1) and every squarefree \(r \ge 1\), \[g_*(r) \;=\; \begin{cases}\dfrac{\mu(r)^2}{g(r)}, & (r, W) = 1,\\[2mm] 0, & (r, W) > 1,\end{cases}\] with \(g\) the multiplicative function of Definition 1.
Proof. Both sides are totally multiplicative in \(r\), so it suffices to compare them at primes. If \(p \mid W\) then \(\gamma_g(p) = 0\), so \(g_*(p) = 0\); total multiplicativity then forces \(g_*(r) = 0\) whenever some prime divisor of \(r\) divides \(W\), i.e. whenever \((r, W) > 1\). If \(p \nmid W\) then \(\gamma_g(p) = p/(p-1)\) and \[g_*(p) \;=\; \frac{\gamma_g(p)}{p - \gamma_g(p)} \;=\; \frac{p/(p-1)}{p - p/(p-1)} \;=\; \frac{1}{p - 2} \;=\; \frac{1}{g(p)} .\] For squarefree \(r\) coprime to \(W\) we have \(\mu(r)^2 = 1\), and multiplying the prime-level identities over the (distinct) primes dividing \(r\) gives \(g_*(r) = 1/g(r)\). ◻
Proof uses: lem_g_star_well_defined
A family of densities \((\gamma_N)_{N \ge 1}\) is called \(W\)-tricked if each \(\gamma_N\) is multiplicative and non-negative with \(\gamma_N(1) = 1\), if \(\gamma_N(p) = 0\) for every prime \(p \mid W\), and if there is a single constant \(c > 0\), independent of \(N\), with \(|\gamma_N(p) - 1| \le c/p\) for every prime \(p \nmid W\). Both \(\gamma_g\) and \(\gamma^{(m)}\) are \(W\)-tricked, with \(c\) absolute. The next two sublemmas supply the two Mertens-deviation constants that a sieve datum (Definition 1) requires.
Sublemma 1 (Uniform upper Mertens deviation \(A_2\)). Let \((\gamma_N)_{N \ge 1}\) be a \(W\)-tricked family with constant \(c\). Then there is \(A_2 > 0\), depending only on \(c\) and in particular independent of \(N\), such that for every \(N\) with \(D_0 \ge 2\) and all reals \(2 \le w \le z\), \[\Delta(\gamma_N; w, z) \;\le\; A_2 .\]
Uses: def_W_trick, def_D_0, lem_mertens_first
Proof. Step 1: compare \(\gamma_N\) with the constant density. For a prime \(p \nmid W\) write \(\gamma_N(p)(\log p)/p = (\log p)/p + (\gamma_N(p) - 1)(\log p)/p\); the second term is at most \(c(\log p)/p^2\) in absolute value, and \(\sum_p (\log p)/p^2\) converges. For \(p \mid W\) the term vanishes altogether, and \(p \mid W\) holds exactly when \(p \le D_0\). Hence there is \(E = E(c)\) with \[\Bigl|\sum_{w \le p \le z}\frac{\gamma_N(p)\log p}{p} \;-\; \sum_{\max(w,\, D_0) \le p \le z}\frac{\log p}{p}\Bigr| \;\le\; E\] for all \(2 \le w \le z\).
Step 2: Mertens’ first theorem. By Lemma 1, discarding the proper prime powers as in the proof of Lemma 1, there is an absolute \(M\) with \(\bigl|\sum_{p \le t}(\log p)/p - \log t\bigr| \le M\) for \(t \ge 2\); hence \(\sum_{a \le p \le z}(\log p)/p \le \log(z/a) + 2M\) for all \(2 \le a \le z\).
Step 3. Apply Step 2 with \(a := \max(w, D_0) \ge w\). Since \(a \ge w\) we have \(\log(z/a) \le \log(z/w)\), so \[\Delta(\gamma_N; w, z) \;\le\; E + \log(z/a) + 2M - \log(z/w) \;\le\; E + 2M .\] Thus \(A_2 := E(c) + 2M\) works, and it does not depend on \(N\), \(w\) or \(z\). (The gain over the two-sided estimate below is that truncating the prime range upward to \(\max(w, D_0)\) can only decrease the sum, so the \(\log D_0\) discrepancy appears with a favourable sign.) ◻
Proof uses: lem_mertens_first, lem_mertens_tau_k
Sublemma 1 (Lower Mertens deviation \(L \ll \log D_0\)). Let \((\gamma_N)_{N \ge 1}\) be a \(W\)-tricked family with constant \(c\). Then there is \(C > 0\), depending only on \(c\), such that for every \(N\) with \(D_0 \ge 2\) there exists a real \(L\) with \[0 \;\le\; L \;\le\; C\log D_0 \qquad\text{and}\qquad -L \;\le\; \Delta(\gamma_N; w, z) \quad\text{for all } 2 \le w \le z .\]
Proof. Retain the notation of the previous proof. Steps 1 and 2 there give, for \(a = \max(w, D_0)\), \[\Delta(\gamma_N; w, z) \;\ge\; -E + \log(z/a) - 2M - \log(z/w) \;=\; -E - 2M - \log\frac{a}{w} .\] Now \(w \le a = \max(w, D_0) \le w\,D_0\) (using \(w \ge 2\) and \(D_0 \ge 2\)), so \(0 \le \log(a/w) \le \log D_0\). Hence \(\Delta(\gamma_N; w, z) \ge -\bigl(\log D_0 + E + 2M\bigr)\). Since \(D_0 \ge 2\) we have \(\log D_0 \ge \log 2 > 0\), so \(L := \log D_0 + E + 2M \le C\log D_0\) with \(C := 1 + (E + 2M)/\log 2\), which depends only on \(c\). Non-negativity of \(L\) is clear. ◻
Proof uses: slem_gg_A2_upper, lem_mertens_first
Sublemma 1 (Singular product for \(\gamma_g\)). There are constants \(C > 0\) and \(N_0\) such that, for every \(N \ge N_0\) with \(D_0 \ge 2\), \[\Bigl|\mathfrak{S}(\gamma_g) \;-\; \frac{\varphi(W)}{W}\Bigr| \;\le\; C\,\frac{\varphi(W)}{W}\cdot\frac{1}{D_0},\] with \(\mathfrak{S}\) as in Notation 1. Equivalently, \(\mathfrak{S}(\gamma_g) = \frac{\varphi(W)}{W}\bigl(1 + O(1/D_0)\bigr)\).
Proof. Split the defining product \(\mathfrak{S}(\gamma_g) = \prod_p (1 - \gamma_g(p)/p)^{-1} (1 - 1/p)\) over \(p \mid W\) and \(p \nmid W\). For \(p \mid W\) one has \(\gamma_g(p) = 0\) and the local factor is \(1 - 1/p\); the product of these over \(p \mid W\) is \(\varphi(W)/W\). For \(p \nmid W\) one has \(\gamma_g(p) = p/(p-1)\), so \[\Bigl(1 - \frac{\gamma_g(p)}{p}\Bigr)^{-1}\Bigl(1 - \frac1p\Bigr) \;=\; \Bigl(1 - \frac{1}{p-1}\Bigr)^{-1}\frac{p-1}{p} \;=\; \frac{(p-1)^2}{p(p-2)} \;=\; 1 + \frac{1}{p(p-2)} .\] Every such \(p\) exceeds \(D_0\), so the tail product satisfies \(1 \le \prod_{p > D_0}\bigl(1 + \frac{1}{p(p-2)}\bigr) \le \exp\bigl(\sum_{p > D_0}\frac{1}{p(p-2)}\bigr) \le \exp(2/D_0)\), using \(\sum_{n > t}n^{-2} \le 2/t\) and \(p(p-2) \ge p^2/3\) for \(p \ge 3\). Since \(\mathrm{e}^{2/D_0} - 1 \le 4/D_0\) for \(D_0 \ge 4\), the tail product is \(1 + O(1/D_0)\), and multiplying by \(\varphi(W)/W\) gives the claim. ◻
The convolution defect and the asymptotic for \(H\)
Sublemma 1 (Size of the convolution defect). Let \(b\) be the convolution defect of Definition 1. For every prime \(p \nmid V\), \[|b(p)| \;\le\; \frac{2A_3}{A_1}\cdot\frac{1}{p^2} .\]
Uses: def_sieve_datum, def_b_defect, def_g_star
Proof. Fix a prime \(p \nmid V\). The structural hypothesis of Definition 1 gives \(|\gamma(p) - 1| \le A_3/p\), and Lemma 1 gives \(p - \gamma(p) \ge A_1 p\). Then \[b(p) \;=\; \frac{\gamma(p)}{p - \gamma(p)} - \frac{1}{p - 1} \;=\; \frac{p\bigl(\gamma(p) - 1\bigr)}{\bigl(p - \gamma(p)\bigr)(p - 1)},\] so, using \(p - 1 \ge p/2\), \[|b(p)| \;\le\; \frac{p\cdot A_3/p}{A_1 p \cdot p/2} \;=\; \frac{2A_3}{A_1 p^2}. \qedhere\] ◻
Proof uses: lem_g_star_well_defined
Definition 1 (The normalised defect \(\tilde b\)). With \(b\) as in Definition 1, set \[\tilde b(e) \;:=\; b(e)\,\frac{\varphi(e)}{e} \qquad (e \ge 1).\] Like \(b\), the function \(\tilde b\) is multiplicative and supported on squarefree integers coprime to \(V\), and \(|\tilde b(e)| \le |b(e)|\) because \(\varphi(e)/e \in [0,1]\).
Uses: def_b_defect, def_totient
Sublemma 1 (Uniform summability of the defect). There is a function \(K = K(A_1, A_3) > 0\), depending only on \(A_1\) and \(A_3\), such that for every sieve datum with those parameters \[\sum_{e \ge 1}|b(e)|\,\tau(e)\,\sqrt{e} \;\le\; K(A_1, A_3) .\]
Uses: def_sieve_datum, def_b_defect, slem_b_bound
Proof. The function \(e \mapsto |b(e)|\tau(e)\sqrt e\) is non-negative, multiplicative, supported on squarefree \(e\), and takes the value \(1\) at \(e = 1\). Hence its sum over all \(e\) equals the Euler product \(\prod_p\bigl(1 + |b(p)|\tau(p)\sqrt p\bigr)\), the primes dividing \(V\) contributing the factor \(1\). For \(p \nmid V\), Sublemma 1 and \(\tau(p) = 2\) give \[|b(p)|\,\tau(p)\sqrt p \;\le\; \frac{4A_3}{A_1}\cdot\frac{1}{p^{3/2}} .\] Since \(\sum_p p^{-3/2}\) converges, \(\log\prod_p(1 + \cdot) \le \sum_p |b(p)|\tau(p)\sqrt p \le (4A_3/A_1)\sum_p p^{-3/2}\), and exponentiating gives a bound depending only on \(A_1\) and \(A_3\). (In particular the sum is finite, so the Euler factorisation used above is legitimate.) ◻
Proof uses: slem_b_bound
Sublemma 1 (Convolution identity for \(h\)). Let \(h\) be as in Definition 1, and let \(a : \mathbb{N}\to\mathbb{R}\) satisfy \(a(n)=\mu(n)^2/\varphi(n)\) for every \(n\). For every squarefree \(d \ge 1\) with \((d, V) = 1\), \[h(d) \;=\; \sum_{e \mid d} b(e)\, a(d/e) .\]
Uses: def_h_conv, def_b_defect, def_g_star, def_totient, def_mobius
Proof. For squarefree \(d\), each divisor \(e \mid d\) is squarefree and coprime to \(d/e\), so the right-hand side is the value at \(d\) of the Dirichlet convolution of two multiplicative functions restricted to squarefree arguments; it therefore factors as \(\prod_{p \mid d}\bigl(b(p) + a(p)\bigr)\), each prime going wholly into \(e\) or into \(d/e\). Since \((d, V) = 1\), every \(p \mid d\) satisfies \(p \nmid V\), so \(a(p) = 1/(p-1)\) and \(b(p) + a(p) = g_*(p)\) by Definition 1. Hence the product equals \(\prod_{p \mid d}g_*(p) = g_*(d) = \mu(d)^2 g_*(d) = h(d)\), using total multiplicativity of \(g_*\) and \(\mu(d)^2 = 1\). ◻
Sublemma 1 (Convolution form of \(H\)). Let \(H\) be as in Definition 1. For every real \(z\), \[H(z) \;=\; \sum_{e \ge 1} b(e) \sum_{\substack{0 < f < z/e\\ (f, Ve) = 1}}\frac{\mu(f)^2}{\varphi(f)} ,\] the \(e\)-sum having only finitely many non-zero terms (those with \(e < z\)).
Uses: def_H_z, def_h_conv, def_b_defect, slem_h_convolution, def_A_m_x
Proof. First, \(h\) is supported on integers coprime to \(V\): this is Lemma 1, whose proof is that for \(p \mid V\) the structural hypothesis gives \(\gamma(p) = 0\), hence \(g_*(p) = 0\), and total multiplicativity of \(g_*\) forces \(h(d) = \mu(d)^2 g_*(d) = 0\) whenever \((d, V) > 1\). Therefore \(H(z) = \sum_{0 < d < z,\ (d, V) = 1}h(d)\), and only squarefree \(d\) contribute.
Now insert Sublemma 1 and exchange the order of the two finite sums, grouping by the divisor \(e\): \[H(z) \;=\; \sum_{\substack{0 < d < z\\ (d, V) = 1,\ d \text{ squarefree}}}\ \sum_{e \mid d} b(e)\,a(d/e) \;=\; \sum_{e \ge 1} b(e) \sum_{\substack{0 < f < z/e\\ (f, Ve) = 1}} a(f) ,\] where \(f = d/e\): the constraints “\(d < z\), \(e \mid d\)” become “\(f < z/e\)”, while “\(d\) squarefree, \((d, V) = 1\)” becomes “\(e\) squarefree coprime to \(V\)” (already enforced by the support of \(b\)) together with “\(f\) squarefree, \((f, Ve) = 1\)” (the condition \((f, e) = 1\) being enforced by \(\mu(f)^2\) inside \(a(f)\) when combined with the squarefreeness of \(d = ef\)). Terms with \(e \ge z\) have empty inner sum. ◻
Proof uses: slem_h_convolution, lem_h_support, def_g_star, def_sieve_datum
Sublemma 1 (Euler-product identification of \(\mathfrak{S}(\gamma)\)). For every sieve datum, \[\mathfrak{S}(\gamma) \;=\; \frac{\varphi(V)}{V}\sum_{e \ge 1}\tilde b(e) ,\] the series converging absolutely. In particular \(\mathfrak{S}(\gamma) > 0\).
Proof. Absolute convergence follows from \(|\tilde b(e)| \le |b(e)| \le |b(e)|\tau(e)\sqrt e\) and Sublemma 1. Since \(\tilde b\) is multiplicative, supported on squarefree integers coprime to \(V\), the sum factors as \(\sum_{e \ge 1}\tilde b(e) = \prod_{p \nmid V}\bigl(1 + \tilde b(p)\bigr)\). For \(p \nmid V\), \[1 + \tilde b(p) \;=\; 1 + \frac{p-1}{p}\Bigl(g_*(p) - \frac{1}{p-1}\Bigr) \;=\; \frac{p-1}{p}\bigl(1 + g_*(p)\bigr) \;=\; \Bigl(1 - \frac1p\Bigr)\Bigl(1 - \frac{\gamma(p)}{p}\Bigr)^{-1},\] using \(1 + g_*(p) = 1 + \gamma(p)/(p - \gamma(p)) = (1 - \gamma(p)/p)^{-1}\). These are exactly the local factors of \(\mathfrak{S}(\gamma)\) (Notation 1) at the primes \(p \nmid V\). At the primes \(p \mid V\) one has \(\gamma(p) = 0\), so the local factor of \(\mathfrak{S}(\gamma)\) is \(1 - 1/p\), and \(\prod_{p \mid V}(1 - 1/p) = \varphi(V)/V\). Multiplying the two contributions gives the identity. Positivity holds because each local factor is positive and the product converges. ◻
Proof uses: slem_b_defect_sum, def_g_star, def_b_tilde
Sublemma 1 (Asymptotic for \(H\)). There are functions \(C_1 = C_1(A_1, A_3) > 0\) and \(C_2 = C_2(A_1, A_3) > 0\) such that, for every sieve datum with parameters \(A_1, A_2, A_3, L\) and every real \(z \ge 2\), \[\bigl|H(z) - \mathfrak{S}(\gamma)\log z\bigr| \;\le\; C_1\,\mathfrak{S}(\gamma)\,(1 + \ell_V) \;+\; C_2\,\tau(V)\,z^{-1/8}\log(2Vz) .\] Both constants depend on \(A_1\) and \(A_3\) only; neither depends on \(A_2\), \(L\), \(\gamma\), \(V\) or \(z\).
Uses: def_H_z, def_sieve_datum, not_singular_series, def_ell_m, slem_H_convolution_form, slem_coprime_density, slem_b_defect_sum, slem_singular_series_btilde, def_b_tilde
Proof. Throughout, all implied constants depend on \(A_1, A_3\) only. Write \(\mathcal{A}(x; m) := \sum_{0 < f < x,\ (f, m) = 1}\mu(f)^2/\varphi(f)\) for the strict truncation appearing in Sublemma 1; it differs from \(A_m(x)\) of Definition 1 by at most one term \(\mu(f)^2/\varphi(f) \ll x^{-1/2}\), so Sublemma 1 applies to \(\mathcal{A}\) with the same error shape.
Preliminaries on \(\tilde b\). From \(|\tilde b(e)| \le |b(e)|\) and Sublemma 1, \[\sum_e |\tilde b(e)|\,\tau(e)\,e^{1/2} = O(1), \quad \sum_e |\tilde b(e)|\,e^{1/4} = O(1), \quad \sum_e |\tilde b(e)|\,(1 + \log e) = O(1),\] the second and third following from the first by \(e^{1/4} \le \tau(e)e^{1/2}\) and \(1 + \log e \le 3\tau(e)e^{1/2}\).
Step 1: main range \(e \le \sqrt z\). For such \(e\) we have \(z/e \ge \sqrt z \ge 2\), so Sublemma 1 applies with \(m = Ve\). Using \((e, V) = 1\), so that \(\varphi(Ve)/(Ve) = (\varphi(V)/V)(\varphi(e)/e)\) and \(\ell_{Ve} = \ell_V + \ell_e\), and \(b(e)\varphi(e)/e = \tilde b(e)\), \[\sum_{\substack{e \le \sqrt z\\ (e, V) = 1}} b(e)\,\mathcal{A}(z/e; Ve) = \frac{\varphi(V)}{V}\sum_{\substack{e \le \sqrt z\\ (e, V) = 1}}\tilde b(e) \bigl(\log z - \log e + \ell_V + \ell_e\bigr) + O\Bigl(\frac{\varphi(V)}{V}\Bigr) + \mathcal{R}(z),\] where the \(O(\varphi(V)/V)\) collects the \(O(\varphi(Ve)/(Ve))\) errors against \(\sum_e |\tilde b(e)| = O(1)\), and \[\mathcal{R}(z) \;:=\; \sum_{\substack{e \le \sqrt z\\ (e, V) = 1}} b(e)\, O\Bigl(\frac{\tau(Ve)\log(2Vez)}{\sqrt{z/e}}\Bigr) \;=\; O\Bigl(\tau(V)\,z^{-1/2}\log(2Vz)\sum_e |b(e)|\tau(e)e^{1/2}\Bigr),\] which is \(O\bigl(\tau(V)z^{-1/2}\log(2Vz)\bigr)\) by Sublemma 1 (using \(\tau(Ve) = \tau(V)\tau(e)\) for \((e, V) = 1\)).
Step 2: complete the \(e\)-sum. Extend the sum to all \(e\) coprime to \(V\). The discarded range \(e > \sqrt z\) is controlled by \(|\log z - \log e + \ell_V + \ell_e| \ll \log(2z) + \ell_V + \log e\) together with \(\sum_{e > \sqrt z}|\tilde b(e)| \le z^{-1/8}\sum_e |\tilde b(e)|e^{1/4} = O(z^{-1/8})\) and the same estimate with an extra \(\log e\); this gives an error \(O\bigl(\frac{\varphi(V)}{V}z^{-1/8}(\log 2z + \ell_V)\bigr)\). Similarly the tail \(e > \sqrt z\) of the true sum of Sublemma 1 is \(O\bigl(\frac{\varphi(V)}{V}z^{-1/8}\log 2z\bigr)\), since \(\mathcal{A}(z/e; Ve) \ll \frac{\varphi(V)}{V}\log(2z)\) trivially.
For the completed sum, split \(\log z - \log e + \ell_V + \ell_e\). By Sublemma 1, \(\frac{\varphi(V)}{V}\sum_{(e, V) = 1}\tilde b(e) = \mathfrak{S}(\gamma)\), so the \(\log z\) and \(\ell_V\) parts produce \(\mathfrak{S}(\gamma)(\log z + \ell_V)\). The remaining part is \[\frac{\varphi(V)}{V}\Bigl|\sum_{(e, V) = 1}\tilde b(e)\bigl(\ell_e - \log e\bigr)\Bigr| \;\ll\; \frac{\varphi(V)}{V}\sum_e |\tilde b(e)|(1 + \log e) \;=\; O\Bigl(\frac{\varphi(V)}{V}\Bigr),\] using \(0 \le \ell_e \le \log e\) for \(e \ge 1\).
Step 3: comparison of \(\varphi(V)/V\) with \(\mathfrak{S}(\gamma)\). By Sublemma 1 and \(\sum_e |\tilde b(e)| = O(1)\) we get \(\mathfrak{S}(\gamma) \ll \varphi(V)/V\); conversely, the local factors at \(p \nmid V\) satisfy \(\log\bigl(1 + \tilde b(p)\bigr) \ge -2A_3/p^2\) by Sublemma 1, so \(\prod_{p \nmid V}(1 + \tilde b(p)) \ge \exp\bigl(-2A_3\sum_p p^{-2}\bigr) \gg 1\) and hence \(\varphi(V)/V \ll \mathfrak{S}(\gamma)\). The two quantities are therefore comparable, with constants depending on \(A_1, A_3\) only, and every \(O(\varphi(V)/V)\) above may be rewritten as \(O(\mathfrak{S}(\gamma))\), hence absorbed into \(O(\mathfrak{S}(\gamma)(1 + \ell_V))\).
Collecting Steps 1–3 and noting \(z^{-1/2} \le z^{-1/8}\) and \(\log(2Vz) \ge 1\) for \(z \ge 2\), \(V \ge 1\), the two error shapes of the statement are obtained. ◻
Partial summation
Sublemma 1 (Cost of the removed primes). There is a function \(C = C(A_3) > 0\) such that, for every sieve datum with parameters \(A_1, A_2, A_3, L\), \[\ell_V \;=\; \sum_{p \mid V}\frac{\log p}{p - 1} \;\le\; 2L \;+\; C(A_3) .\]
Proof. Fix \(z \ge \max\{p : p \mid V\}\) (and \(z \ge 2\)). Since \(\gamma(p) = 0\) for \(p \mid V\), \[\sum_{2 \le p \le z}\frac{\gamma(p)\log p}{p} \;=\; \sum_{\substack{p \le z\\ p \nmid V}}\frac{\log p}{p} \;+\; \sum_{\substack{p \le z\\ p \nmid V}}\frac{(\gamma(p) - 1)\log p}{p} .\] By the structural hypothesis the second sum is at most \(A_3\sum_p (\log p)/p^2 = O_{A_3}(1)\) in absolute value, the series converging. For the first, Mertens’ first theorem \(\sum_{p \le z}(\log p)/p = \log z + O(1)\) (Lemma 1, after discarding the proper prime powers) gives \[\sum_{2 \le p \le z}\frac{\gamma(p)\log p}{p} \;=\; \log z \;-\; \sum_{p \mid V}\frac{\log p}{p} \;+\; O_{A_3}(1).\] The lower half of the log-sum hypothesis at \(w = 2\) reads \(\sum_{2 \le p \le z}\gamma(p)(\log p)/p \ge \log(z/2) - L\). Subtracting yields \(\sum_{p \mid V}(\log p)/p \le L + O_{A_3}(1)\). Finally \(p/(p-1) \le 2\) for every prime \(p\), so \(\ell_V \le 2\sum_{p \mid V}(\log p)/p \le 2L + O_{A_3}(1)\). ◻
Lemma 1 (Partial summation). There are functions \(C_1 = C_1(A_1, A_2, A_3) > 0\) and \(C_2 = C_2(A_1, A_3) > 0\) such that the following holds for every sieve datum with parameters \(A_1, A_2, A_3, L\), every continuously differentiable \(G : \mathbb{R} \to \mathbb{R}\), and every real \(z \ge 2\): \[\Bigl|\sum_{0 < d < z}\mu(d)^2 g_*(d)\,G\Bigl(\frac{\log d}{\log z}\Bigr) \;-\; \mathfrak{S}(\gamma)\,(\log z)\int_0^1 G(x)\,dx\Bigr| \;\le\; C_1\,\mathfrak{S}(\gamma)\,L\,G_{\max} \;+\; C_2\,\frac{\tau(V)\log(2Vz)}{\log z}\,G_{\max},\] where \(G_{\max} = \sup_{t \in [0,1]}\bigl(|G(t)| + |G'(t)|\bigr)\).
Uses: def_sieve_datum, def_g_star, def_h_conv, def_H_z, not_singular_series, slem_H_asymptotic, slem_removed_primes
Proof. Write \(S(z) := \sum_{0 < d < z}h(d)G(\log d/\log z)\) with \(h(d) = \mu(d)^2 g_*(d)\).
Step 1: Abel summation. The function \(H\) of Definition 1 is the summatory function of \(h\); taking \(\psi(t) := G(\log t/\log z)\), so that \(\psi'(t) = G'(\log t/\log z)/(t\log z)\) and \(\psi(z) = G(1)\), Abel summation gives the exact identity \[S(z) \;=\; G(1)\,H(z) \;-\; \int_1^z \frac{G'(\log t/\log z)}{t\log z}\,H(t)\,dt .\] (The integrand is continuous off the integers and \(H\) is a bounded monotone step function on \([1, z]\), so the integral exists.)
Step 2: separate main term and error. Write \(H(t) = \mathfrak{S}(\gamma)\log t + r(t)\). For the main term, substitute \(x := \log t/\log z\), so \(dx = dt/(t\log z)\) and \(\log t = x\log z\): \[\mathfrak{S}(\gamma)\int_1^z\frac{G'(\log t/\log z)}{t\log z}\log t\,dt \;=\; \mathfrak{S}(\gamma)(\log z)\int_0^1 x\,G'(x)\,dx .\] Integration by parts gives \(\int_0^1 xG'(x)\,dx = G(1) - \int_0^1 G(x)\,dx\), so together with the boundary contribution \(G(1)\mathfrak{S}(\gamma)\log z\) the main terms combine to exactly \(\mathfrak{S}(\gamma)(\log z)\int_0^1 G(x)\,dx\). Consequently \[\Bigl|S(z) - \mathfrak{S}(\gamma)(\log z)\int_0^1 G\Bigr| \;\le\; G_{\max}\Bigl(|r(z)| + \frac{1}{\log z}\int_1^z\frac{|r(t)|}{t}\,dt\Bigr),\] using \(|G(1)| \le G_{\max}\) and \(|G'| \le G_{\max}\) on \([0,1]\).
Step 3: bound the error functional. By Sublemma 1, \[|r(t)| \;\le\; C_1'\,\mathfrak{S}(\gamma)(1 + \ell_V) \;+\; C_2'\,\tau(V)t^{-1/8}\log(2Vt) \qquad (2 \le t \le z),\] with \(C_1', C_2'\) depending on \(A_1, A_3\); on \([1, 2]\) one has \(|H(t)| \le 1\) and \(\mathfrak{S}(\gamma)\log t \le \mathfrak{S}(\gamma)\log 2\), so the same shape holds after enlarging the constants. The first piece integrates to \(O\bigl(\mathfrak{S}(\gamma)(1 + \ell_V)\bigr)\) (since \(\frac{1}{\log z}\int_1^z dt/t = 1\)), and the second to \(O\bigl(\tau(V)\log(2V)/\log z\bigr)\), because \(\int_1^\infty t^{-9/8}\log(2Vt)\,dt = O(\log 2V)\) is a convergent integral. The boundary term \(|r(z)|\) is dominated by the same two shapes, using \(z^{-1/8} \le 1/\log z\) for \(z \ge 2\).
Step 4: collapse \(\ell_V\) to \(L\). By Sublemma 1, \(1 + \ell_V \le 1 + 2L + C(A_3)\). Absorbing the additive constants (the \(A_2\)-dependence entering through the parameter \(L\) supplied by the log-sum hypothesis) turns \(\mathfrak{S}(\gamma)(1 + \ell_V)\) into \(C_1(A_1, A_2, A_3)\mathfrak{S}(\gamma)L\) plus a term of the second shape. Multiplying through by \(G_{\max}\) gives the stated bound. ◻
Proof uses: slem_H_asymptotic, slem_removed_primes, def_ell_m, def_H_z
Relating the transformed weight systems
The sieve analysis uses four systems of variables: the profile \(F\), the divisor weights \(\lambda(d_1, \dotsc, d_k)\), the \(y\)-variables \(y(r_1, \dotsc, r_k)\) (Definition 1), and, for each \(m \in \{1, \dotsc, k\}\), the \(m\)-th marginal variables \(y^{(m)}(r_1, \dotsc, r_k)\) (Definition 1). The \(S_1(\lambda)\) analysis is carried out in the \(y\)-variables and the \(S_2^{(m)}(\lambda)\) analysis in the \(y^{(m)}\)-variables; both are evaluated in terms of \(F\). This section relates them: the \(y\)-variables of the \(F\)-derived sieve weight are evaluated in closed form, and \(y^{(m)}\) is expressed in terms of \(y\) up to an error of relative size \(1/D_0\).
Throughout, \(k \ge 2\) is fixed, \(\lambda : \mathbb{N}^k \to \mathbb{R}\) is a function of finite support, \(y\) denotes the \(y\)-transform of \(\lambda\) (Definition 1) and \(y^{(m)}\) the \(m\)-th marginal transform (Definition 1). The parameters \(D_0\), \(W\) and \(R\) are those of Definitions 1, 1 and 1, and “permissible support” always means Definition 1.
Finiteness of the marginal supremum
Lemma 1 (\(y^{(m)}_{\max}\) is a finite maximum). Let \(m \in \{1, \dotsc, k\}\) and let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be a function of finite support. Then \(y^{(m)}_{\max} < \infty\); in fact the supremum is attained as a maximum on a finite subset of \(\mathbb{N}^k\).
Uses: def_ym_vars, def_y_max
Proof. Suppose \(y^{(m)}(r_1, \dotsc, r_k) \ne 0\). By Definition 1 the defining \(\mathbf{d}\)-sum has a nonzero summand, so there is a tuple \(\mathbf{d} = (d_1, \dotsc, d_k) \in \operatorname{supp}(\lambda)\) with \(r_i \mid d_i\) for every \(i\); in particular \(r_i \le d_i\) for every \(i\). Since \(\lambda\) has finite support, the integer \(M := \max\{d_i : \mathbf{d} \in \operatorname{supp}(\lambda),\ 1 \le i \le k\}\) is well defined and finite, and every \(\mathbf{r}\) with \(y^{(m)}(\mathbf{r}) \ne 0\) lies in the finite set \(\{1, \dotsc, M\}^k\). Hence \(y^{(m)}_{\max} = \max\{|y^{(m)}(\mathbf{r})| : \mathbf{r} \in \{1, \dotsc, M\}^k\}\) is a maximum over a finite set, and is finite. ◻
The \(y\)-variables of the \(F\)-derived weight
Lemma 1 (The transform of the \(F\)-derived weight is exactly \(y_0\)). Let \(R > 0\), let \(W \ge 1\), and let \(F : \mathbb{R}^k \to \mathbb{R}\) be arbitrary. Let \(y_0\) be the \(F\)-derived \(y\)-weight of Definition 1 and let \(\lambda_0\) be the associated sieve weight of Definition 1. Then the \(y\)-transform of \(\lambda_0\) (Definition 1) is \(y_0\): \[y_{\lambda_0}(r_1, \dotsc, r_k) \;=\; y_0(r_1, \dotsc, r_k) \qquad \text{for every } (r_1, \dotsc, r_k) \in \mathbb{N}^k .\]
Proof. By Definition 1 the weight \(\lambda_0\) is the Möbius transform \(\lambda_{y_0}\) of \(y_0\) (Definition 1). The round-trip identity Lemma 1 therefore gives \[y_{\lambda_0}(\mathbf{r}) \;=\; y_{\lambda_{y_0}}(\mathbf{r}) \;=\; \begin{cases} y_0(\mathbf{r}) & \text{if $r_i$ is squarefree for every $i$,}\\ 0 & \text{otherwise.} \end{cases}\] It remains to check that \(y_0(\mathbf{r}) = 0\) whenever some \(r_i\) fails to be squarefree. In that case \(\prod_j r_j\) is not squarefree, so \(\mathbf{r}\) does not lie in the permissible support of Definition 1; since \(y_0\) is by construction supported there (Definition 1), \(y_0(\mathbf{r}) = 0\). Both sides therefore agree at every \(\mathbf{r}\). ◻
Proof uses: def_lambda_from_F, lem_y_from_lambda, def_lambda_from_y, def_adm_support
Lemma 1 (Piecewise closed form of \(y\) for the \(F\)-derived sieve weight). Let \(k \ge 2\), let \(R > 0\), let \(W \ge 1\), and let \(F : \mathbb{R}^k \to \mathbb{R}\) be a smooth function with \(\operatorname{supp} F \subseteq \mathcal{R}_k\) (Definition 1). Let \(\lambda_0\) be the \(F\)-derived sieve weight of Definition 1 and let \(y\) be its \(y\)-transform (Definition 1). Then, for every \((r_1, \dotsc, r_k) \in \mathbb{N}^k\), \[y(r_1, \dotsc, r_k) \;=\; \begin{cases} F\bigl(\tfrac{\log r_1}{\log R}, \dotsc, \tfrac{\log r_k}{\log R}\bigr) & \text{if $\prod_{i=1}^k r_i$ is squarefree and $\gcd\bigl(\prod_{i=1}^k r_i,\, W\bigr) = 1$,}\\[2pt] 0 & \text{otherwise.} \end{cases}\]
Proof. By Lemma 1, \(y = y_0\), and by Definition 1, \[y_0(\mathbf{r}) \;=\; F\Bigl(\tfrac{\log r_1}{\log R}, \dotsc, \tfrac{\log r_k}{\log R}\Bigr) \cdot \mathbf{1}\bigl[\mathbf{r} \text{ lies in the permissible support}\bigr].\] The permissible support consists of the tuples satisfying the three conditions \(\prod_i r_i \le R\), \(\gcd(\prod_i r_i, W) = 1\), and \(\prod_i r_i\) squarefree (Definition 1). Two of these are exactly the conditions in the assertion, so it suffices to prove that the truncation condition is automatic: precisely, that \[F\Bigl(\tfrac{\log r_1}{\log R}, \dotsc, \tfrac{\log r_k}{\log R}\Bigr) \ne 0 \quad \text{and} \quad \textstyle\prod_i r_i \text{ squarefree} \qquad \Longrightarrow \qquad \textstyle\prod_i r_i \le R .\] Assume the left-hand side and write \(x_i := \log r_i / \log R\). Squarefreeness of \(\prod_i r_i\) forces \(r_i \ge 1\) for every \(i\), so \(\log r_i \ge 0\). Since \(F\) is nonzero at \(\mathbf{x} = (x_1, \dotsc, x_k)\) and \(\operatorname{supp} F \subseteq \mathcal{R}_k\), we have \(\mathbf{x} \in \mathcal{R}_k\), i.e. \(x_i \ge 0\) for every \(i\) and \(\sum_i x_i \le 1\).
If \(\log R > 0\), then \(\sum_i x_i \le 1\) reads \(\sum_i \log r_i \le \log R\), that is, \(\log(\prod_i r_i) \le \log R\); exponentiating gives \(\prod_i r_i \le R\), as required.
If instead \(\log R \le 0\) (that is, \(0 < R \le 1\)), then each \(x_i = \log r_i/\log R\) is a quotient of a non-negative number by a non-positive one, hence \(x_i \le 0\); combined with \(x_i \ge 0\) this gives \(\mathbf{x} = \mathbf{0}\). But \(F(\mathbf{0}) = 0\): since \(k \ge 1\), the origin is a limit of points outside \(\mathcal{R}_k\) (move along a single coordinate ray into negative values), \(F\) vanishes identically off \(\mathcal{R}_k\), and \(F\) is continuous, so \(F(\mathbf{0}) = 0\) by continuity. This contradicts \(F(\mathbf{x}) \ne 0\), so the case is vacuous. In either case the implication holds, and the piecewise formula follows. ◻
Proof uses: lem_y_from_lambda_0, def_y_0, def_adm_support
Definition 1 (\(C^1\)-norm \(F_{\max}\) on \([0,1]^k\)). For a \(C^1\) function \(F : \mathbb{R}^k \to \mathbb{R}\), set \[F_{\max} \;:=\; \sup_{t \in [0,1]^k}\Bigl(|F(t)| \;+\; \sum_{i=1}^k \Bigl|\frac{\partial F}{\partial t_i}(t)\Bigr|\Bigr).\] The supremum is over the compact unit cube, on which the displayed function is continuous, so it is finite; it is non-negative because the quantity inside the supremum is non-negative.
Expressing \(y^{(m)}\) through \(y\)
Lemma 1 (Substitute the inversion formula for \(\lambda\) into \(y^{(m)}\)). Let \(\lambda\) have permissible support, let \(m \in \{1, \dotsc, k\}\), and let \(\mathbf{r} \in \mathbb{N}^k\) satisfy \(r_m = 1\). Then \[y^{(m)}(r_1, \dotsc, r_k) \;=\; \Bigl(\prod_{i=1}^k \mu(r_i)\, g(r_i)\Bigr) \sum_{\substack{\mathbf{d} \in \mathbb{N}^k \\ r_i \mid d_i\ \forall i,\ d_m = 1}} \Bigl(\prod_{i=1}^k \frac{\mu(d_i)\, d_i}{\varphi(d_i)}\Bigr) \sum_{\substack{\mathbf{a} \in \mathbb{N}^k \\ d_i \mid a_i\ \forall i}} \frac{y(a_1, \dotsc, a_k)}{\prod_{i=1}^k \varphi(a_i)} ,\] both sums running over tuples of positive integers.
Proof. By Definition 1, \[y^{(m)}(r_1, \dotsc, r_k) \;=\; \Bigl(\prod_{i=1}^k \mu(r_i)\, g(r_i)\Bigr) \sum_{\substack{\mathbf{d}\\ r_i \mid d_i\ \forall i,\ d_m = 1}} \frac{\lambda(d_1, \dotsc, d_k)}{\prod_{i=1}^k \varphi(d_i)} ,\] a finite sum because \(\lambda\) has finite support; moreover every contributing \(\mathbf{d}\) has \(\prod_i d_i\) squarefree, hence positive squarefree coordinates, by Definition 1. On such tuples Lemma 1 identifies \(\lambda\) with the Möbius transform of its own \(y\)-variables (Definition 1): \[\lambda(\mathbf{d}) \;=\; \Bigl(\prod_{i=1}^k \mu(d_i)\, d_i\Bigr) \sum_{\substack{\mathbf{a}\\ d_i \mid a_i\ \forall i}} \frac{y(a_1, \dotsc, a_k)}{\prod_{i=1}^k \varphi(a_i)} .\] Dividing by \(\prod_i \varphi(d_i)\) gives \[\frac{\lambda(\mathbf{d})}{\prod_{i=1}^k \varphi(d_i)} \;=\; \Bigl(\prod_{i=1}^k \frac{\mu(d_i)\, d_i}{\varphi(d_i)}\Bigr) \sum_{\substack{\mathbf{a}\\ d_i \mid a_i\ \forall i}} \frac{y(a_1, \dotsc, a_k)}{\prod_{i=1}^k \varphi(a_i)} ,\] and inserting this into the previous display yields the asserted formula. (The hypothesis \(r_m = 1\) is not used in the derivation; it is recorded because \(y^{(m)}\) vanishes otherwise, by Lemma 1.) ◻
Proof uses: lem_lambda_from_y, def_lambda_from_y, lem_ym_support
Lemma 1 (Interchange the \(\mathbf{d}\)- and \(\mathbf{a}\)-summations). Let \(\lambda\) have permissible support, let \(m \in \{1, \dotsc, k\}\), and let \(\mathbf{r} \in \mathbb{N}^k\) satisfy \(r_i \ge 1\) for every \(i\) and \(r_m = 1\). Then \[y^{(m)}(r_1, \dotsc, r_k) \;=\; \Bigl(\prod_{i=1}^k \mu(r_i)\, g(r_i)\Bigr) \sum_{\substack{\mathbf{a} \in \mathbb{N}^k \\ r_i \mid a_i\ \forall i}} \frac{y(a_1, \dotsc, a_k)}{\prod_{i=1}^k \varphi(a_i)} \sum_{\substack{\mathbf{d} \in \mathbb{N}^k \\ r_i \mid d_i \mid a_i\ \forall i,\ d_m = 1}} \prod_{i=1}^k \frac{\mu(d_i)\, d_i}{\varphi(d_i)} .\]
Proof. The iterated sum of Lemma 1 is supported on the set of pairs \((\mathbf{d}, \mathbf{a})\) with \(r_i \mid d_i \mid a_i\) for every \(i\) and \(d_m = 1\): the outer summation imposes \(r_i \mid d_i\) and \(d_m = 1\), and the inner one imposes \(d_i \mid a_i\), which together force \(r_i \mid a_i\). Both summations are finite, since \(y\) has finite support (Definition 1) and, for fixed \(\mathbf{a}\), only the finitely many divisor tuples \(\mathbf{d}\) of \(\mathbf{a}\) contribute. We may therefore interchange the order of summation: grouping first by \(\mathbf{a}\), the inner \(\mathbf{d}\)-sum runs over \(\{\mathbf{d} : r_i \mid d_i \mid a_i\ \forall i,\ d_m = 1\}\). Multiplying by the outer prefactor \(\prod_i \mu(r_i) g(r_i)\), which depends only on \(\mathbf{r}\), gives the asserted identity. ◻
Proof uses: lem_ym_substitute_y
Lemma 1 (Evaluation of the inner divisor sum). Let \(a\) be a positive squarefree integer and let \(r\) be a positive divisor of \(a\). Then \[\sum_{\substack{d \mid a\\ r \mid d}} \frac{\mu(d)\, d}{\varphi(d)} \;=\; \frac{\mu(a)\, r}{\varphi(a)} .\]
Uses: def_mobius, def_totient
Proof. Write \(a = r e\) with \(e := a/r\). Since \(a\) is squarefree, so are \(r\) and \(e\), and \((r, e) = 1\).
Every \(d\) with \(r \mid d \mid a\) is uniquely of the form \(d = r f\) with \(f \mid e\), and then \((r, f) = 1\), so by multiplicativity of \(\mu\) and of \(\varphi\) on coprime products, \(\mu(d) = \mu(r)\mu(f)\) and \(\varphi(d) = \varphi(r)\varphi(f)\). Substituting, \[\sum_{\substack{d \mid a\\ r \mid d}} \frac{\mu(d)\, d}{\varphi(d)} \;=\; \frac{\mu(r)\, r}{\varphi(r)} \sum_{f \mid e} \frac{\mu(f)\, f}{\varphi(f)} .\] The summand \(f \mapsto \mu(f) f/\varphi(f)\) is multiplicative, so for squarefree \(e\) the divisor sum factors over the prime divisors of \(e\): \[\sum_{f \mid e} \frac{\mu(f)\, f}{\varphi(f)} \;=\; \prod_{p \mid e}\Bigl(1 + \frac{\mu(p)\, p}{\varphi(p)}\Bigr) \;=\; \prod_{p \mid e}\Bigl(1 - \frac{p}{p-1}\Bigr) \;=\; \prod_{p \mid e} \frac{-1}{p-1} \;=\; \frac{(-1)^{\omega(e)}}{\prod_{p \mid e}(p-1)} \;=\; \frac{\mu(e)}{\varphi(e)} ,\] using \(\mu(e) = (-1)^{\omega(e)}\) and \(\varphi(e) = \prod_{p \mid e}(p-1)\) for squarefree \(e\). Combining the two displays, \[\sum_{\substack{d \mid a\\ r \mid d}} \frac{\mu(d)\, d}{\varphi(d)} \;=\; \frac{\mu(r)\,\mu(e)\, r}{\varphi(r)\,\varphi(e)} \;=\; \frac{\mu(a)\, r}{\varphi(a)} ,\] the last step again by multiplicativity of \(\mu\) and \(\varphi\) on the coprime product \(a = re\). ◻
Lemma 1 (\(y^{(m)}\) in terms of \(y\), intermediate form). Let \(\lambda\) have permissible support, let \(m \in \{1, \dotsc, k\}\), and let \(\mathbf{r} \in \mathbb{N}^k\) satisfy \(r_i \ge 1\) for every \(i\) and \(r_m = 1\). Then \[y^{(m)}(r_1, \dotsc, r_k) \;=\; \Bigl(\prod_{i=1}^k \mu(r_i)\, g(r_i)\Bigr) \sum_{\substack{\mathbf{a} \in \mathbb{N}^k\\ r_i \mid a_i\ \forall i}} \frac{y(a_1, \dotsc, a_k)}{\prod_{i=1}^k \varphi(a_i)} \prod_{i \ne m} \frac{\mu(a_i)\, r_i}{\varphi(a_i)} .\]
Proof. Start from Lemma 1 and evaluate, for each fixed \(\mathbf{a}\), the inner sum \[\Sigma(\mathbf{a}) \;:=\; \sum_{\substack{\mathbf{d}\\ r_i \mid d_i \mid a_i\ \forall i,\ d_m = 1}} \prod_{i=1}^k \frac{\mu(d_i)\, d_i}{\varphi(d_i)} .\] It suffices to treat those \(\mathbf{a}\) with \(y(\mathbf{a}) \ne 0\); by Lemma 1 such an \(\mathbf{a}\) has \(\prod_i a_i\) squarefree, so every \(a_i\) is positive and squarefree. The constraints on \(\mathbf{d}\) are coordinatewise, so \(\Sigma(\mathbf{a})\) factors as a product over \(i\). The \(m\)-th factor is \(1\): the constraint \(d_m = 1\) leaves the single term \(\mu(1)\cdot 1/\varphi(1) = 1\), and \(r_m = 1\) is compatible with it. For \(i \ne m\) the \(i\)-th factor is \(\sum_{r_i \mid d_i \mid a_i} \mu(d_i)d_i/\varphi(d_i)\), which equals \(\mu(a_i) r_i/\varphi(a_i)\) by Lemma 1 applied with \(a = a_i\) and \(r = r_i\) (legitimate since \(a_i\) is squarefree and \(r_i \mid a_i\)). Hence \(\Sigma(\mathbf{a}) = \prod_{i \ne m} \mu(a_i) r_i / \varphi(a_i)\), and substituting this into Lemma 1 gives the assertion. ◻
Proof uses: lem_ym_swap, lem_ym_eval_d_sum, lem_y_permissible_support
The diagonal decomposition of \(y^{(m)}\)
Definition 1 (The \(y^{(m)}\)-summand \(T_{m, \mathbf{r}}\)). Let \(m \in \{1, \dotsc, k\}\) and \(\mathbf{r} \in \mathbb{N}^k\). For \(\mathbf{a} \in \mathbb{N}^k\) set \[T_{m, \mathbf{r}}(\mathbf{a}) \;:=\; \Bigl(\prod_{i=1}^k \mu(r_i)\, g(r_i)\Bigr) \cdot \frac{y(a_1, \dotsc, a_k)}{\prod_{i=1}^k \varphi(a_i)} \cdot \prod_{i \ne m} \frac{\mu(a_i)\, r_i}{\varphi(a_i)} ,\] with the convention that the whole expression is read as \(0\) when some \(a_i = 0\).
Uses: def_y_from_lambda, def_mobius, def_g_func, def_totient
Definition 1 (The free-coordinate marginal \(Y^{(m)}\)). Let \(m \in \{1, \dotsc, k\}\) and \(\mathbf{r} \in \mathbb{N}^k\). Set \[Y^{(m)}(r_1, \dotsc, r_k) \;:=\; \sum_{a \ge 1} \frac{y(r_1, \dotsc, r_{m-1},\, a,\, r_{m+1}, \dotsc, r_k)}{\varphi(a)} ,\] the sum running over the single free coordinate in the \(m\)-th slot. Only finitely many terms are nonzero, because \(y\) has finite support.
Uses: def_y_from_lambda, def_totient
Lemma 1 (Diagonal and off-diagonal decomposition of \(y^{(m)}\)). Let \(\lambda\) have permissible support, let \(m \in \{1, \dotsc, k\}\), and let \(\mathbf{r} \in \mathbb{N}^k\) be such that every \(r_i\) is positive and squarefree and \(r_m = 1\). Then \[y^{(m)}(r_1, \dotsc, r_k) \;=\; \Bigl(\prod_{i \ne m} \frac{g(r_i)\, r_i}{\varphi(r_i)^2}\Bigr) \, Y^{(m)}(r_1, \dotsc, r_k) \;+\; \sum_{\substack{\mathbf{a} \in \mathbb{N}^k :\ r_i \mid a_i\ \forall i \\ \exists\, j \ne m,\ a_j \ne r_j}} T_{m, \mathbf{r}}(\mathbf{a}) .\]
Proof. By Lemma 1 and Definition 1, \[y^{(m)}(\mathbf{r}) \;=\; \sum_{\substack{\mathbf{a}\\ r_i \mid a_i\ \forall i}} T_{m, \mathbf{r}}(\mathbf{a}) ,\] the sum being finite because \(y\) has finite support and \(T_{m, \mathbf{r}}\) vanishes wherever \(y\) does. Split the index set \(\{\mathbf{a} : r_i \mid a_i\ \forall i\}\) into the diagonal part \(\mathcal{D} := \{\mathbf{a} : r_i \mid a_i\ \forall i,\ a_i = r_i\ \forall i \ne m\}\) and its complement, the off-diagonal part \(\mathcal{O} := \{\mathbf{a} : r_i \mid a_i\ \forall i,\ \exists j \ne m,\ a_j \ne r_j\}\). These are disjoint with union the whole index set, so the sum splits accordingly, and the second piece is the displayed off-diagonal sum.
It remains to evaluate the diagonal piece. Since \(r_m = 1\) divides every integer, the tuples in \(\mathcal{D}\) are exactly those obtained from \(\mathbf{r}\) by replacing the \(m\)-th coordinate by an arbitrary \(a \ge 0\), and distinct \(a\) give distinct tuples. For such a tuple, using \(\varphi(a_i) = \varphi(r_i)\) and \(\mu(a_i) = \mu(r_i)\) for \(i \ne m\), \[T_{m, \mathbf{r}}(\mathbf{a}) \;=\; \Bigl(\mu(r_m) g(r_m) \prod_{i \ne m} \mu(r_i) g(r_i)\Bigr) \cdot \frac{y(\mathbf{a})}{\varphi(a)\prod_{i \ne m}\varphi(r_i)} \cdot \prod_{i \ne m} \frac{\mu(r_i)\, r_i}{\varphi(r_i)} .\] The factor \(\mu(r_m) g(r_m)\) equals \(1\) because \(r_m = 1\), and \(\mu(r_i)^2 = 1\) for each \(i \ne m\) because \(r_i\) is squarefree. Collecting the \(i \ne m\) factors, \[T_{m, \mathbf{r}}(\mathbf{a}) \;=\; \Bigl(\prod_{i \ne m} \frac{g(r_i)\, r_i}{\varphi(r_i)^2}\Bigr) \cdot \frac{y(\mathbf{a})}{\varphi(a)} .\] Summing over \(a \ge 1\) — the term \(a = 0\) contributes \(0\) by the convention of Definition 1 — and recognising Definition 1 gives the diagonal contribution \(\bigl(\prod_{i \ne m} g(r_i) r_i/\varphi(r_i)^2\bigr)\, Y^{(m)}(\mathbf{r})\), as asserted. ◻
Proof uses: lem_ym_intermediate, def_mobius
Lemma 1 (Approximation of the diagonal prefactor). There is a constant \(C = C(k) \ge 0\) with the following property. Let \(D_0 \ge 2\) be real and let \(\mathbf{r} \in \mathbb{N}^k\) be such that every \(r_i\) is positive and squarefree and every prime factor of every \(r_i\) exceeds \(D_0\). Then \[\Bigl| \prod_{i=1}^k \frac{g(r_i)\, r_i}{\varphi(r_i)^2} \;-\; 1 \Bigr| \;\le\; \frac{C}{D_0} .\]
Uses: def_g_func, def_totient, def_D_0
Proof. Fix \(i\). Since \(r_i\) is squarefree, both \(g\) and \(\varphi\) factor over the distinct prime divisors of \(r_i\), and \[\frac{g(r_i)\, r_i}{\varphi(r_i)^2} \;=\; \prod_{p \mid r_i} \frac{(p-2)\, p}{(p-1)^2} \;=\; \prod_{p \mid r_i} \Bigl(1 - \frac{1}{(p-1)^2}\Bigr) ,\] because \((p-2)p = (p-1)^2 - 1\). Hence \[\prod_{i=1}^k \frac{g(r_i)\, r_i}{\varphi(r_i)^2} \;=\; \prod_{j} (1 - x_j), \qquad x_j := \frac{1}{(p_j - 1)^2},\] where \(p_1, p_2, \dotsc\) enumerates, with multiplicity, the prime divisors of \(r_1, \dotsc, r_k\). Every such prime exceeds \(D_0 \ge 2\), so \(p_j \ge 3\) and \(0 \le x_j \le 1/4 \le 1\). For numbers \(x_j \in [0,1]\) one has the elementary bound \(\bigl|\prod_j (1 - x_j) - 1\bigr| \le \sum_j x_j\), proved by induction on the number of factors: if \(P\) is a partial product then \(0 \le P \le 1\), so \(|P(1-x) - 1| \le |P - 1| + P x \le |P-1| + x\).
It therefore suffices to bound \(\sum_j x_j\). For a single index \(i\) the primes \(p \mid r_i\) are distinct and all exceed \(D_0\), and \(p - 1 \ge p/2\) for every prime \(p\), so \[\sum_{p \mid r_i} \frac{1}{(p-1)^2} \;\le\; 4 \sum_{n > D_0} \frac{1}{n^2} \;\le\; \frac{4}{\lfloor D_0 \rfloor} \;\le\; \frac{8}{D_0},\] using the integral comparison \(\sum_{n > M} n^{-2} \le \int_{M}^{\infty} t^{-2}\,dt = 1/M\) at the integer \(M = \lfloor D_0 \rfloor\), and \(\lfloor D_0 \rfloor \ge D_0/2\) for \(D_0 \ge 1\). Summing over the \(k\) coordinates gives \(\sum_j x_j \le 8k/D_0\), so the lemma holds with \(C := 8k\). ◻
Lemma 1 (Bound for the free-coordinate marginal). There is a constant \(C = C(k) > 0\) with the following property. Let \(W \ge 1\) and let \(R \ge 1\) be real with \(p \le R\) for every prime \(p \mid W\). Let \(\lambda\) have permissible support, with \(y\)-variables \(y\). Then, for every \(m \in \{1, \dotsc, k\}\) and every \(\mathbf{r} \in \mathbb{N}^k\), \[\bigl| Y^{(m)}(r_1, \dotsc, r_k) \bigr| \;\le\; C\, \frac{y_{\max}\, \varphi(W)\, (\log R + 1)}{W} .\]
Proof. Write \(\mathbf{r}^{(a)}\) for the tuple obtained from \(\mathbf{r}\) by putting \(a\) in the \(m\)-th slot. By Lemma 1, \(y(\mathbf{r}^{(a)}) = 0\) unless \(\prod_i r^{(a)}_i\) is squarefree, coprime to \(W\), and at most \(R\); in particular \(a\) must lie in \[\mathcal{S} \;:=\; \{ a \in \mathbb{N} : a \le R,\ a \text{ squarefree},\ (a, W) = 1 \} .\] Hence, by the triangle inequality and \(|y| \le y_{\max}\), \[\bigl|Y^{(m)}(\mathbf{r})\bigr| \;\le\; y_{\max} \sum_{a \in \mathcal{S}} \frac{1}{\varphi(a)} .\] It remains to bound \(A := \sum_{a \in \mathcal{S}} 1/\varphi(a)\). For squarefree \(a\) one has the divisor-sum identity \(a/\varphi(a) = \sum_{d \mid a} \mu(d)^2/\varphi(d)\), so \(1/\varphi(a) = a^{-1}\sum_{d \mid a} \mu(d)^2/\varphi(d)\). Writing \(a = de\) and dropping the constraints linking \(d\) and \(e\) (all terms being non-negative), \[A \;\le\; \Bigl(\sum_{\substack{d \ge 1\\ d\ \text{squarefree},\ (d, W) = 1}} \frac{\mu(d)^2}{\varphi(d)\, d}\Bigr) \cdot \Bigl(\sum_{\substack{e \le R\\ (e, W) = 1}} \frac{1}{e}\Bigr) .\] The first factor is bounded by an absolute constant \(K\): extending it to all squarefree \(d \ge 1\) and factoring the resulting Euler product gives \(\prod_p \bigl(1 + 1/(p(p-1))\bigr) < \infty\). The second factor is the Mertens-type coprime harmonic sum, which by Lemma 1 is \(O\bigl((\varphi(W)/W)(\log R + 1)\bigr)\) under the stated hypothesis that every prime factor of \(W\) is at most \(R\). Multiplying the two bounds gives \(A \le C\, \varphi(W)(\log R + 1)/W\) for a constant \(C = C(k) > 0\), and the lemma follows. ◻
Proof uses: lem_y_permissible_support, lem_mertens_reciprocal_W, def_mobius
Lemma 1 (Off-diagonal contribution: \(a_j \ne r_j\) for some \(j \ne m\)). There is a constant \(C = C(k) > 0\) with the following property. Let \(\vartheta \in (0,1)\) and \(\delta \in (0, \vartheta/2)\). Then there is \(N_0\) such that for every \(N \ge N_0\), every \(y\) of permissible support, every \(m \in \{1, \dotsc, k\}\), and every \(\mathbf{r} \in \mathbb{N}^k\) with \(r_i \ge 1\) for every \(i\) and \(r_m = 1\), \[\Bigl| \sum_{\substack{\mathbf{a} \in \mathbb{N}^k :\ r_i \mid a_i\ \forall i \\ \exists\, j \ne m,\ a_j \ne r_j}} T_{m, \mathbf{r}}(\mathbf{a}) \Bigr| \;\le\; C\, \frac{y_{\max}\, \varphi(W)\, \log R}{W\, D_0} .\]
Uses: def_ym_summand, def_y_max, def_adm_support, def_W_trick, def_D_0, def_R
Proof. If \(k \le 1\) the off-diagonal index set is empty and there is nothing to prove, so assume \(k \ge 2\).
Step 1: a \(y\)-free majorant. On the permissible support every coordinate \(a_i\) is positive, squarefree and coprime to \(W\), and \(\prod_i a_i \le R\). Since \(|y(\mathbf{a})| \le y_{\max}\) and \(|\mu| \le 1\), Definition 1 gives, for every permissible \(\mathbf{a}\), \[\bigl|T_{m, \mathbf{r}}(\mathbf{a})\bigr| \;\le\; y_{\max} \cdot \Bigl|\prod_{i=1}^k \mu(r_i) g(r_i)\Bigr| \cdot \frac{1}{\prod_{i=1}^k \varphi(a_i)} \cdot \prod_{i \ne m} \frac{r_i}{\varphi(a_i)} .\] Summing over the off-diagonal set, the whole off-diagonal sum is at most \(y_{\max}\, K\), where \(K\) denotes the sum of the right-hand majorant (without the factor \(y_{\max}\)) over all permissible off-diagonal \(\mathbf{a}\).
Step 2: reduction to one-dimensional sums. Every contributing \(\mathbf{a}\) has \(r_i \mid a_i\); write \(a_i = r_i b_i\) with \(b_i \ge 1\). On the permissible support each \(a_i\) is squarefree, so \((r_i, b_i) = 1\) and \(\varphi(a_i) = \varphi(r_i)\varphi(b_i)\); moreover each \(b_i\) is itself squarefree, coprime to \(W\), and at most \(R\), that is, \(b_i \in \mathcal{S}\), where \(\mathcal{S} := \{n \le R : n \text{ squarefree},\ (n, W) = 1\}\). In these variables the majorant factors as \[\underbrace{\Bigl|\prod_{i} \mu(r_i) g(r_i)\Bigr| \cdot \frac{1}{\prod_i \varphi(r_i)} \cdot \prod_{i \ne m}\frac{r_i}{\varphi(r_i)}}_{=: \, \Pi(\mathbf{r})} \;\cdot\; \frac{1}{\varphi(b_m)} \prod_{i \ne m} \frac{1}{\varphi(b_i)^2} .\] Now \(\Pi(\mathbf{r}) \le 1\): indeed \(|\mu| \le 1\), the \(m\)-th coordinate contributes \(g(1)/\varphi(1) = 1\) since \(r_m = 1\), and for each \(i \ne m\) the squarefreeness of \(r_i\) gives \(g(r_i)\, r_i \le \varphi(r_i)^2\), this being the inequality \(\prod_{p \mid r_i}\bigl(1 - (p-1)^{-2}\bigr) \le 1\). Being off-diagonal means \(a_j \ne r_j\), i.e. \(b_j \ne 1\), for at least one \(j \ne m\); bounding the number of such \(j\) by \(k - 1\) and summing each coordinate independently, \[K \;\le\; (k-1)\, A \cdot T \cdot B^{\,k-2}, \qquad A := \sum_{n \in \mathcal{S}} \frac{1}{\varphi(n)}, \quad B := \sum_{n \in \mathcal{S}} \frac{1}{\varphi(n)^2}, \quad T := \sum_{\substack{n \in \mathcal{S}\\ n \ne 1}} \frac{1}{\varphi(n)^2} ,\] where \(A\) accounts for the free \(m\)-th coordinate, \(T\) for the distinguished offending coordinate \(j\), and \(B\) for each of the remaining \(k - 2\) coordinates.
Step 3: the three sums. First, \(B \le \sum_{n \ge 1}\mu(n)^2/\varphi(n)^2 =: C_B\), which is finite by Lemma 1 and absolute. Second, every \(n \in \mathcal{S}\) with \(n \ne 1\) is squarefree and coprime to \(W\), hence has all of its prime factors exceeding \(D_0\) (Definition 1); choosing one such prime factor \(p\) and writing \(n = p n'\) with \(n' \in \mathcal{S}\) coprime to \(p\), we have \(\varphi(n) = (p-1)\varphi(n')\), so, over-counting each \(n\) at most once per prime factor, \[T \;\le\; \Bigl(\sum_{p > D_0} \frac{1}{(p-1)^2}\Bigr) \cdot B \;\le\; \frac{8}{D_0} \cdot C_B ,\] the bound on the prime sum being the one established in the proof of Lemma 1. Third, \(A \le C_A\, \varphi(W)(\log R + 1)/W\) exactly as in the proof of Lemma 1, via Lemma 1.
Step 4: assembly. Combining, \[K \;\le\; (k-1)\, C_A\, C_B^{\,k-1} \cdot \frac{8}{D_0} \cdot \frac{\varphi(W)(\log R + 1)}{W} .\] Since \(R = N^{\vartheta/2 - \delta}\) with \(\vartheta/2 - \delta > 0\), we have \(\log R \ge 1\) for all \(N\) beyond a threshold \(N_0\) depending only on \(\vartheta\) and \(\delta\); for such \(N\), \(\log R + 1 \le 2\log R\). Multiplying by \(y_{\max}\) from Step 1 gives the assertion with \(C := 16(k-1)\, C_A\, C_B^{\,k-1}\). ◻
Lemma 1 (\(y^{(m)}\) from \(y\)). There is a constant \(C = C(k) > 0\) with the following property. Let \(\vartheta \in (0,1)\) and \(\delta \in (0, \vartheta/2)\). Then there is \(N_0\) such that for every \(N \ge N_0\), every \(\lambda\) of permissible support, every \(m \in \{1, \dotsc, k\}\), and every \(\mathbf{r} \in \mathbb{N}^k\) with \(r_m = 1\) and every \(r_i\) squarefree, \[\bigl| y^{(m)}(r_1, \dotsc, r_k) \;-\; Y^{(m)}(r_1, \dotsc, r_k) \bigr| \;\le\; C\, \frac{y_{\max}\, \varphi(W)\, \log R}{W\, D_0} .\]
Proof. Squarefreeness of each \(r_i\) gives \(r_i \ge 1\), so Lemma 1 applies and, writing \(c(m, \mathbf{r}) := \prod_{i \ne m} g(r_i) r_i/\varphi(r_i)^2\), \[y^{(m)}(\mathbf{r}) - Y^{(m)}(\mathbf{r}) \;=\; \bigl(c(m, \mathbf{r}) - 1\bigr)\, Y^{(m)}(\mathbf{r}) \;+\; \sum_{\substack{\mathbf{a} :\ r_i \mid a_i\ \forall i\\ \exists\, j \ne m,\ a_j \ne r_j}} T_{m, \mathbf{r}}(\mathbf{a}) .\] The second term is bounded by Lemma 1, applied to the \(y\)-variables of \(\lambda\), which have permissible support by Lemma 1; this gives \(C_1\, y_{\max}\varphi(W)\log R/(W D_0)\) for all \(N\) beyond a threshold. We bound the first term.
If some \(r_j\) with \(j \ne m\) is not coprime to \(W\), then every tuple obtained from \(\mathbf{r}\) by varying the \(m\)-th coordinate has coordinate product not coprime to \(W\), so \(y\) vanishes at all of them by Lemma 1; hence \(Y^{(m)}(\mathbf{r}) = 0\) and the first term vanishes.
Otherwise every \(r_i\) is squarefree and coprime to \(W\), so by Definition 1 every prime factor of every \(r_i\) exceeds \(D_0\). Applying Lemma 1 to the tuple \(\mathbf{r}\) — whose \(m\)-th factor equals \(1\), so that the full product and the product over \(i \ne m\) agree — gives \(|c(m, \mathbf{r}) - 1| \le C_2/D_0\) once \(D_0 \ge 2\), which holds for all large \(N\). Combined with Lemma 1, \[\bigl|(c(m, \mathbf{r}) - 1)\, Y^{(m)}(\mathbf{r})\bigr| \;\le\; \frac{C_2}{D_0} \cdot C_3\, \frac{y_{\max}\,\varphi(W)(\log R + 1)}{W} \;\le\; 2 C_2 C_3\, \frac{y_{\max}\,\varphi(W)\log R}{W D_0} ,\] using \(\log R + 1 \le 2\log R\), valid for all \(N\) beyond a threshold depending only on \(\vartheta\) and \(\delta\). Adding the two contributions, and taking \(N_0\) to be the largest of the finitely many thresholds used, gives the claim with \(C := C_1 + 2C_2C_3\). ◻
The smooth case
Lemma 1 (Substitute smooth \(y\) into Lemma 1). There is a constant \(C = C(k) \ge 0\) with the following property. Let \(\vartheta \in (0,1)\), \(\delta \in (0, \vartheta/2)\), and let \(F : \mathbb{R}^k \to \mathbb{R}\) be smooth with \(\operatorname{supp} F \subseteq \mathcal{R}_k\). Then there is \(N_0\) such that for every \(N \ge N_0\), every \(\lambda\) of permissible support whose \(y\)-variables satisfy \(y_{\max} \le F_{\max}\), every \(m \in \{1, \dotsc, k\}\), and every \(\mathbf{r} \in \mathbb{N}^k\) with \(r_m = 1\) and every \(r_i\) squarefree, \[\bigl| y^{(m)}(r_1, \dotsc, r_k) \;-\; Y^{(m)}(r_1, \dotsc, r_k) \bigr| \;\le\; C\, \frac{F_{\max}\, \varphi(W)\, \log R}{W\, D_0} .\]
Uses: lem_ym_from_y, def_F_smooth_max, def_ym_vars, def_ym_marginal, def_y_max, def_simplex, def_adm_support, def_W_trick, def_D_0, def_R
Proof. Apply Lemma 1, which bounds the left-hand side by \(C\, y_{\max}\varphi(W)\log R/(W D_0)\) for all \(N\) beyond a threshold. For all such \(N\) we have \(\log R > 0\), and \(\varphi(W)/W > 0\) always, so that bound is a non-decreasing function of \(y_{\max}\). Replacing \(y_{\max}\) by the larger quantity \(F_{\max}\) using the hypothesis \(y_{\max} \le F_{\max}\) yields the stated bound with the same constant. ◻
Lemma 1 (Maximal modulus of the transformed coefficients). Let \(k \ge 1\), let \(R \ge 2\), and let \(y : \mathbb{N}^k \to \mathbb{R}\) be finitely supported with permissible support at level \(R\) and modulus \(W\) (Definition 1). Then the transformed system \(\lambda_y\) of Definition 1 satisfies \[\|\lambda_y\|_{\max} \;\le\; e^{\,1 + 3k}\, \|y\|_{\max}\, (\log R)^{k}.\]
Proof. Fix \(\mathbf{d}\). Expanding \(\lambda_y(\mathbf d)\) and bounding each summand by \(\|y\|_{\max}\) leaves the divisor sum \(\sum_{\mathbf r} \prod_i \mu(r_i)^2/\varphi(r_i)\) taken over those \(\mathbf r\) with \(d_i \mid r_i\) and \(\prod_i r_i \le R\). Together with the prefactor \(\prod_i \mu(d_i)\,d_i\) this is exactly the quantity bounded in Lemma 1, taken at modulus \(W\) and at the present \(\mathbf{d}\). Each coordinate factor is at most \(\sum_{r \le R} \mu(r)^2/\varphi(r) \le e\,(1 + \log R)\), and the outer product over the \(k\) coordinates contributes at most \(e^{3k}\) after absorbing the prefactor \(\prod_i \mu(d_i) d_i\) against the coprimality constraints. Multiplying the \(k\) coordinate bounds and using \(1 + \log R \le e \log R\) for \(R \ge 2\) gives the stated bound. ◻
Proof uses: lem_restricted_reciprocal_sum
Sieve manipulations: the reduction of \(S_1(\lambda)\)
For the first moment \(S_1(\lambda)\), we pass from the coefficients \(\lambda(d_1, \dotsc, d_k)\) to the \(y\)-variables. We expand the square defining \(w_n\) and interchange the order of summation (Lemma 1); we evaluate the resulting inner count by the Chinese remainder theorem and bound the aggregate error (Lemmas 1–1); we split \(1/[d_i, e_i]\) across the two coordinates, detect the residual cross-coprimality by Möbius inversion, and restrict the resulting off-diagonal variables (Lemmas 1–1); we substitute the inversion formula expressing \(\lambda\) in terms of \(y\), and show that all but the diagonal term \(s_{i, j} \equiv 1\) is negligible (Lemmas 1–1); and finally we insert the \(F\)-derived weight, drop the last coprimality constraint, and convert the discrete sum to an integral by partial summation (Lemmas 1–1). The final statement, Lemma 1, is the asymptotic \[S_1(\lambda_0) \;=\; \frac{\varphi(W)^k\, N\, (\log R)^k}{W^{k+1}}\, I_k(F) \;+\; O_k\Bigl(\frac{F_{\max}^2\, \varphi(W)^k\, N\, (\log R)^k}{W^{k+1}\, D_0}\Bigr).\]
Throughout, \(k \ge 2\) is fixed, \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) is admissible (Definition 1), \(\vartheta \in (0, 1)\) and \(\delta \in (0, \vartheta/2)\) are fixed real parameters, \(R = N^{\vartheta/2 - \delta}\) (Definition 1), \(D_0 = \log\log\log N\) (Definition 1), \(W = \prod_{p \le D_0} p\) (Definition 1), and \(v_0\) is a valid residue class (Lemma 1). No level-of-distribution hypothesis is used in this section: \(\vartheta\) enters only through the truncation \(R\), and the arithmetic input is elementary (Mertens sums and partial summation). Bombieri–Vinogradov is needed only for the second moment.
Auxiliary objects
Definition 1 (The lcm modulus \(q(\mathbf{d}, \mathbf{e})\)). Let \(\mathbf{d} = (d_1, \dotsc, d_k)\) and \(\mathbf{e} = (e_1, \dotsc, e_k)\) be \(k\)-tuples of positive integers. Define \[q(\mathbf{d}, \mathbf{e}) \;:=\; W \prod_{i=1}^k [d_i, e_i].\]
Uses: def_W_trick, not_bold_tuple
Definition 1 (The CRT remainder \(\varrho(\mathbf{d}, \mathbf{e})\)). For \(k\)-tuples \(\mathbf{d}, \mathbf{e}\) of positive integers, set \[\varrho(\mathbf{d}, \mathbf{e}) \;:=\; \sum_{\substack{N < n \le 2N\\ n \equiv v_0 \pmod W\\ [d_i, e_i] \mid n + h_i\ \forall i}} 1 \;-\; \frac{N}{q(\mathbf{d}, \mathbf{e})}\, \mathbf{1}\bigl[W, [d_1, e_1], \dotsc, [d_k, e_k] \text{ are pairwise coprime}\bigr].\]
Uses: def_q_lcm, def_W_trick, not_bold_tuple
Notation 1 (Restricted sum \(\sideset{}{^*}\sum\) over the \(s_{i, j}\)). Fix positive integers \(u_1, \dotsc, u_k\). For a function \(f\) of the variables \((s_{i, j})_{1 \le i \ne j \le k}\) we write \[\sideset{}{^*}\sum_{s_{1,2}, \dotsc, s_{k, k-1}} f\bigl((s_{i, j})_{i \ne j}\bigr) \;:=\; \sum_{\substack{s_{i, j} \ge 1,\ 1 \le i \ne j \le k\\ (\ast_1),\,(\ast_2),\,(\ast_3),\,(\ast_4)}} f\bigl((s_{i, j})_{i \ne j}\bigr),\] where the four restrictions are \[\begin{aligned} (\ast_1)\ &\ (s_{i, j},\, u_i) = 1 && \text{for all } i \ne j,\\ (\ast_2)\ &\ (s_{i, j},\, u_j) = 1 && \text{for all } i \ne j,\\ (\ast_3)\ &\ (s_{i, j},\, s_{i, a}) = 1 && \text{for all } a \notin \{i, j\},\\ (\ast_4)\ &\ (s_{i, j},\, s_{b, j}) = 1 && \text{for all } b \notin \{i, j\}. \end{aligned}\]
Definition 1 (The exponent tuple \(\mathbf{a}\)). Given positive integers \(u_1, \dotsc, u_k\) and \((s_{i, j})_{1 \le i \ne j \le k}\), set \[a_j \;:=\; u_j \prod_{i \ne j} s_{j, i} \qquad (1 \le j \le k), \qquad \mathbf{a} := (a_1, \dotsc, a_k).\]
Definition 1 (The exponent tuple \(\mathbf{b}\)). With the same data, set \[b_j \;:=\; u_j \prod_{i \ne j} s_{i, j} \qquad (1 \le j \le k), \qquad \mathbf{b} := (b_1, \dotsc, b_k).\] Note the transposition of the indices of \(s\) relative to Definition 1: in \(a_j\) the first index of \(s\) is fixed to \(j\), in \(b_j\) the second.
Uses: def_bold_a
Expansion and the Chinese remainder theorem
Lemma 1 (Expansion of \(S_1(\lambda)\)). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) have finite support. Then \[S_1(\lambda) \;=\; \sum_{\mathbf{d}}\, \sum_{\mathbf{e}}\, \lambda(\mathbf{d})\, \lambda(\mathbf{e}) \sum_{\substack{N < n \le 2N\\ n \equiv v_0 \pmod W\\ [d_i, e_i] \mid n + h_i\ \forall i}} 1 .\]
Uses: def_S1, def_weight, def_W_trick, not_bold_tuple, not_standing
Proof. We perform three rearrangements.
Step 1: expand the square. By Definition 1, \[w_n \;=\; \Bigl(\sum_{d_i \mid n + h_i\ \forall i} \lambda(\mathbf{d})\Bigr)^{\!2} \;=\; \sum_{d_i \mid n + h_i\ \forall i}\ \sum_{e_i \mid n + h_i\ \forall i} \lambda(\mathbf{d})\, \lambda(\mathbf{e}),\] the second equality being the distributive law for the product of two finite sums (both inner sums are over divisors of the fixed positive integers \(n + h_i\), hence finite).
Step 2: substitute into \(S_1\). By Definition 1, \[S_1(\lambda) \;=\; \sum_{\substack{N < n \le 2N\\ n \equiv v_0 \pmod W}} \sum_{d_i \mid n + h_i\ \forall i}\ \sum_{e_i \mid n + h_i\ \forall i} \lambda(\mathbf{d})\, \lambda(\mathbf{e}).\]
Step 3: interchange. All three sums are finite: the outer one is over a finite arithmetic progression in \((N, 2N]\), and the inner ones are supported in the finite support of \(\lambda\). Interchanging, and merging the two divisibilities \(d_i \mid n + h_i\) and \(e_i \mid n + h_i\) into the single condition \([d_i, e_i] \mid n + h_i\), gives the assertion. ◻
Lemma 1 (CRT evaluation of the inner count). Let \(\lambda\) have \(R\)-permissible support (Definition 1), let \(v_0\) be a valid residue class, and let \(\mathbf{d}, \mathbf{e}\) be \(k\)-tuples of positive integers with \(\lambda(\mathbf{d}) \ne 0\) and \(\lambda(\mathbf{e}) \ne 0\). If \(N\) is large enough that \(|h_i - h_j| \le D_0\) for all \(i \ne j\), then \[\sum_{\substack{N < n \le 2N\\ n \equiv v_0 \pmod W\\ [d_i, e_i] \mid n + h_i\ \forall i}} 1 \;=\; \begin{cases} \dfrac{N}{q(\mathbf{d}, \mathbf{e})} + O(1) & \text{if } W, [d_1, e_1], \dotsc, [d_k, e_k] \text{ are pairwise coprime},\\[1.2ex] 0 & \text{otherwise,} \end{cases}\] with an absolute implied constant.
Proof. Write \(q := q(\mathbf{d}, \mathbf{e})\). The summation imposes the system of congruences \(n \equiv v_0 \pmod W\) and \(n \equiv -h_i \pmod{[d_i, e_i]}\) for \(i = 1, \dotsc, k\).
Coprime case. If the \(k+1\) moduli \(W, [d_1, e_1], \dotsc, [d_k, e_k]\) are pairwise coprime, the Chinese remainder theorem gives a unique solution modulo their product, which is exactly \(q\). The interval \((N, 2N]\) has length \(N\), so it contains \(N/q + O(1)\) integers in a fixed residue class modulo \(q\); the error is at most \(1\) in absolute value.
Non-coprime case. Suppose the moduli are not pairwise coprime, so some prime \(p\) divides two of them. If \(p \mid W\) and \(p \mid [d_i, e_i]\) for some \(i\), then \(p \le D_0\) by Definition 1, while \([d_i, e_i]\) divides \(\prod_\ell d_\ell e_\ell\), which is coprime to \(W\) by \(R\)-permissibility of the support (Definition 1); this is a contradiction, so this case does not occur. Otherwise \(p \mid [d_i, e_i]\) and \(p \mid [d_j, e_j]\) for some \(i \ne j\). Any \(n\) counted by the sum satisfies \(p \mid n + h_i\) and \(p \mid n + h_j\), hence \(p \mid h_i - h_j\). As \(p \nmid W\) we have \(p > D_0\), whereas \(0 < |h_i - h_j| \le D_0\) for \(N\) large; so \(h_i - h_j\) is a nonzero integer of absolute value smaller than \(p\), and \(p \mid h_i - h_j\) is impossible. The set of counted \(n\) is therefore empty and the sum is \(0\). ◻
Lemma 1 (Aggregate CRT lookup error). Let \(\delta \in (0, \vartheta/2)\) and let \(\lambda\) have \(R\)-permissible support. Then, for all sufficiently large \(N\), with \(\varrho\) as in Definition 1, \[\Bigl|\sum_{\mathbf{d}} \sum_{\mathbf{e}} \lambda(\mathbf{d})\, \lambda(\mathbf{e})\, \varrho(\mathbf{d}, \mathbf{e})\Bigr| \;\le\; \lambda_{\max}^2 \sum_{\mathbf{d}} \sum_{\mathbf{e}} |\varrho(\mathbf{d}, \mathbf{e})| \;\ll_k\; \lambda_{\max}^2\, R^2\, (\log R)^{2k} \;\ll_k\; y_{\max}^2\, R^2\, (\log R)^{4k},\] the sums being over the (finite) support of \(\lambda\).
Uses: def_adm_support, def_R, def_varrho_summand, def_y_from_lambda, def_y_max, lem_S1_CRT, not_bold_tuple
Proof. The first inequality is the triangle inequality together with \(|\lambda(\mathbf{d})\lambda(\mathbf{e})| \le \lambda_{\max}^2\).
Step 1: each remainder is bounded. By Lemma 1, for every pair \((\mathbf{d}, \mathbf{e})\) in the support of \(\lambda\) we have \(|\varrho(\mathbf{d}, \mathbf{e})| \le C_\varrho\) for an absolute constant \(C_\varrho\): in the pairwise coprime case \(\varrho\) is the \(O(1)\) discrepancy of the count, and in the remaining case both terms defining \(\varrho\) vanish.
Step 2: count the tuples in the support. By \(R\)-permissibility every \(\mathbf{d}\) in the support satisfies \(d_i \ge 1\) for all \(i\) and \(\prod_i d_i \le R\). The number of such tuples is at most \[\#\Bigl\{\mathbf{d} \in \mathbb{N}^k : \prod_i d_i \le R\Bigr\} \;\le\; \Bigl\lfloor R \Bigr\rfloor \cdot \Bigl(\sum_{1 \le a \le \lfloor R\rfloor} \frac1a\Bigr)^{\!k} \;\ll_k\; R\,(\log R)^k .\] The middle bound holds because on the counting set one has \(1 \le R \prod_i d_i^{-1}\), so the cardinality is at most \(R \sum_{\mathbf{d}} \prod_i d_i^{-1}\) with each \(d_i\) running over \([1, \lfloor R\rfloor]\), and this factorises into \(k\) harmonic sums; the last step uses \(\sum_{a \le x} 1/a \le 1 + \log x\).
Step 3: assemble. Squaring Step 2 bounds the number of pairs by \(\ll_k R^2 (\log R)^{2k}\); multiplying by \(C_\varrho \lambda_{\max}^2\) gives the second inequality.
Step 4: pass from \(\lambda_{\max}\) to \(y_{\max}\). By Lemma 1, \(\lambda_{\max} \ll_k y_{\max}(\log R)^k\), so \(\lambda_{\max}^2 \ll_k y_{\max}^2 (\log R)^{2k}\); substituting gives the last bound. ◻
Proof uses: lem_lambda_max_bound
Lemma 1 (Negligibility of the CRT lookup error). Let \(\vartheta \in (0, 1)\), \(\delta \in (0, \vartheta/2)\) and \(Y \ge 0\). Then, for all sufficiently large \(N\), \[Y^2\, R^2\, (\log R)^{4k} \;\ll_k\; \frac{Y^2\, \varphi(W)^k\, N\, (\log R)^k}{W^{k+1}\, D_0}.\]
Uses: def_R, def_D_0, def_W_trick, def_totient
Proof. Both sides carry the factor \(Y^2 \ge 0\), so we may divide it out.
Step 1: polynomial saving. By Definition 1, \(R^2 = N^{\vartheta - 2\delta}\), and \(\vartheta - 2\delta < \vartheta < 1\). Put \(\eta := 1 - (\vartheta - 2\delta) > 0\), so \(R^2 = N^{1 - \eta}\) and, using \(\log R \le \log N\), \[R^2 (\log R)^{4k} \;\le\; N^{1-\eta} (\log N)^{4k}.\]
Step 2: the \(W\)-factor is harmless. By Lemma 1, \(W \ll (\log\log N)^2 \le \log N\) for large \(N\), so \(W^{k+1} \le (\log N)^{2k+2}\); combined with \(\varphi(W) \ge 1\) this gives \(\varphi(W)^k (\log N)^{2k+2}/W^{k+1} \ge 1\).
Step 3: combine. Multiplying the bound of Step 1 by the quantity of Step 2, \[R^2 (\log R)^{4k} \;\le\; N^{1-\eta}\,(\log N)^{6k+2}\,\frac{\varphi(W)^k}{W^{k+1}} \;=\; N\,\frac{\varphi(W)^k}{W^{k+1}} \cdot N^{-\eta}(\log N)^{6k+2}.\] Since \(\log R = (\vartheta/2 - \delta)\log N\) and \(D_0 = \log\log\log N\), the target factor \((\log R)^k/D_0\) is \(\gg_k (\log N)^k / \log\log\log N\), whereas \(N^{-\eta}(\log N)^{6k+2} \to 0\) as \(N \to \infty\) because \(\eta > 0\) dominates any fixed power of \(\log N\). Hence \(N^{-\eta}(\log N)^{6k+2} \le (\log R)^k/D_0\) for large \(N\), which is the claim. ◻
Proof uses: lem_W_size
Lemma 1 (Main term of \(S_1(\lambda)\) after the CRT step). Let \(k \ge 2\), let \(\mathcal{H}\) be admissible, let \(\vartheta \in (0, 1)\) and \(\delta \in (0, \vartheta/2)\). Then, for all sufficiently large \(N\), for every \(\lambda\) with \(R\)-permissible support and every valid \(v_0\), \[S_1(\lambda) \;=\; \frac{N}{W} \sideset{}{'}\sum_{\mathbf{d}, \mathbf{e}} \frac{\lambda(\mathbf{d})\, \lambda(\mathbf{e})}{\prod_{i=1}^k [d_i, e_i]} \;+\; O_k\Bigl(\frac{y_{\max}^2\, \varphi(W)^k\, N\, (\log R)^k}{W^{k+1}\, D_0}\Bigr),\] where \(\sideset{}{'}\sum\) denotes summation restricted to pairs \((\mathbf{d}, \mathbf{e})\) of tuples of positive integers for which \(W, [d_1, e_1], \dotsc, [d_k, e_k]\) are pairwise coprime.
Uses: def_S1, def_admissible, def_adm_support, def_R, def_D_0, def_W_trick, def_y_from_lambda, def_y_max, lem_exists_v0, not_bold_tuple, not_standing
Proof. Step 1. By Lemma 1, \(S_1(\lambda) = \sum_{\mathbf{d}, \mathbf{e}} \lambda(\mathbf{d})\lambda(\mathbf{e})\, T(\mathbf{d}, \mathbf{e})\), where \(T(\mathbf{d}, \mathbf{e})\) is the inner count.
Step 2. Admissibility makes \(h_1, \dotsc, h_k\) distinct, so for \(N\) large enough \(|h_i - h_j| \le D_0\) for all \(i \ne j\), and Lemma 1 applies to every pair in the support. By Definition 1 this says precisely \[T(\mathbf{d}, \mathbf{e}) \;=\; \frac{N}{q(\mathbf{d}, \mathbf{e})}\, \mathbf{1}\bigl[W, [d_1, e_1], \dotsc, [d_k, e_k]\ \text{pairwise coprime}\bigr] \;+\; \varrho(\mathbf{d}, \mathbf{e}).\] Substituting into Step 1 and using \(N/q(\mathbf{d}, \mathbf{e}) = (N/W)\prod_i [d_i, e_i]^{-1}\), \[S_1(\lambda) \;=\; \frac{N}{W} \sideset{}{'}\sum_{\mathbf{d}, \mathbf{e}} \frac{\lambda(\mathbf{d})\lambda(\mathbf{e})}{\prod_i [d_i, e_i]} \;+\; \sum_{\mathbf{d}, \mathbf{e}} \lambda(\mathbf{d})\lambda(\mathbf{e})\, \varrho(\mathbf{d}, \mathbf{e}).\]
Step 3. By Lemma 1 the second sum is \(\ll_k y_{\max}^2 R^2 (\log R)^{4k}\).
Step 4. By Lemma 1 applied with \(Y = y_{\max}\), this is \(\ll_k y_{\max}^2 \varphi(W)^k N (\log R)^k / (W^{k+1} D_0)\), which is the stated error term. ◻
Decoupling the lcm and the Möbius inversion
Lemma 1 (Removing the automatic constraints). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be such that \(\lambda(\mathbf{t}) \ne 0\) implies that every \(t_i\) is positive, \(\prod_i t_i\) is squarefree, and \((\prod_i t_i, W) = 1\). Then, for every function \(f\) of pairs of tuples, \[\sideset{}{'}\sum_{\mathbf{d}, \mathbf{e}} \lambda(\mathbf{d})\, \lambda(\mathbf{e})\, f(\mathbf{d}, \mathbf{e}) \;=\; \sum_{\substack{\mathbf{d}, \mathbf{e}\\ (d_i, e_j) = 1\ \forall\, i \ne j}} \lambda(\mathbf{d})\, \lambda(\mathbf{e})\, f(\mathbf{d}, \mathbf{e}),\] where \(\sideset{}{'}\sum\) is the restricted summation of Lemma 1.
Proof. It suffices to show that, on the support of \(\lambda(\mathbf{d})\lambda(\mathbf{e})\), the pairwise coprimality of \(W, [d_1, e_1], \dotsc, [d_k, e_k]\) is equivalent to \((d_i, e_j) = 1\) for all \(i \ne j\); both sides then have the same summands.
Step 1: the intra-coprimalities are automatic. If \(\lambda(\mathbf{d}) \ne 0\) then \(\prod_i d_i\) is squarefree. If a prime \(p\) divided both \(d_i\) and \(d_j\) with \(i \ne j\), then \(p^2 \mid \prod_\ell d_\ell\), a contradiction. Hence \((d_i, d_j) = 1\) for \(i \ne j\), and likewise \((e_i, e_j) = 1\) for \(i \ne j\).
Step 2: coprimality to \(W\) is automatic. From \((\prod_i d_i, W) = 1\) and \((\prod_i e_i, W) = 1\), and since \([d_i, e_i]\) divides \(\prod_\ell d_\ell e_\ell\), we get \(([d_i, e_i], W) = 1\) for every \(i\).
Step 3: the residual condition. Pairwise coprimality of the \(k+1\) moduli unwinds into \(([d_i, e_i], W) = 1\) for all \(i\) (Step 2) together with \(([d_i, e_i], [d_j, e_j]) = 1\) for \(i \ne j\). A prime \(p\) dividing both \([d_i, e_i]\) and \([d_j, e_j]\) divides one of \(d_i, e_i\) and one of \(d_j, e_j\); by Step 1 the possibilities \(p \mid d_i, d_j\) and \(p \mid e_i, e_j\) are excluded, leaving \(p \mid (d_i, e_j)\) or \(p \mid (d_j, e_i)\). Thus the only non-automatic constraint is \((d_i, e_j) = 1\) for all \(i \ne j\), and conversely that constraint together with Steps 1–2 gives the full pairwise coprimality. ◻
Lemma 1 (Decoupling the outer lcm product). Let \(\lambda : \mathbb{N}^k \to \mathbb{R}\) be such that \(\lambda(\mathbf{t}) \ne 0\) implies that every \(t_i\) is positive and squarefree and that \(t_1, \dotsc, t_k\) are pairwise coprime. Then \[\sum_{\substack{\mathbf{d}, \mathbf{e}\\ (d_i, e_j) = 1\ \forall\, i \ne j}} \frac{\lambda(\mathbf{d})\, \lambda(\mathbf{e})}{\prod_{i=1}^k [d_i, e_i]} \;=\; \sum_{\substack{\mathbf{d}, \mathbf{e}\\ (d_i, e_j) = 1\ \forall\, i \ne j}} \Bigl(\ \sum_{\substack{u_1, \dotsc, u_k\\ u_i \mid (d_i, e_i)\ \forall i}} \prod_{i=1}^k \varphi(u_i)\Bigr) \frac{\lambda(\mathbf{d})\, \lambda(\mathbf{e})}{\prod_{i=1}^k d_i e_i}.\]
Proof. We may work summand by summand, and we may assume \(\lambda(\mathbf{d})\lambda(\mathbf{e}) \ne 0\), since otherwise both summands vanish. Then every \(d_i\) and every \(e_i\) is positive and squarefree, so Lemma 1 gives, for each \(i\), \[\frac{1}{[d_i, e_i]} \;=\; \frac{1}{d_i e_i} \sum_{u_i \mid (d_i, e_i)} \varphi(u_i).\] Taking the product over \(i = 1, \dotsc, k\) and expanding the product of \(k\) finite sums into a single sum over the tuple \((u_1, \dotsc, u_k)\) with \(u_i \mid (d_i, e_i)\), \[\frac{1}{\prod_{i=1}^k [d_i, e_i]} \;=\; \frac{1}{\prod_{i=1}^k d_i e_i} \sum_{\substack{u_1, \dotsc, u_k\\ u_i \mid (d_i, e_i)\ \forall i}} \prod_{i=1}^k \varphi(u_i).\] Multiplying by \(\lambda(\mathbf{d})\lambda(\mathbf{e})\) and summing over the pairs satisfying the cross-coprimality guard gives the identity; the guard is a condition on \((\mathbf{d}, \mathbf{e})\) alone and is untouched by the manipulation, since the \(\mathbf{u}\)-sum is internal to the summand. ◻
Lemma 1 (Möbius inversion of the constraint \((d_i, e_j) = 1\)). For every finite family of coefficients \(\lambda\) and every pair \((\mathbf{d}, \mathbf{e})\), \[\sum_{\substack{\mathbf{d}, \mathbf{e}\\ (d_i, e_j) = 1\ \forall\, i \ne j}} \Bigl(\ \sum_{\substack{u_1, \dotsc, u_k\\ u_i \mid (d_i, e_i)}} \prod_i \varphi(u_i)\Bigr) \frac{\lambda(\mathbf{d})\lambda(\mathbf{e})}{\prod_i d_i e_i} \;=\; \Sigma_{\mathrm{full}},\] where \(\Sigma_{\mathrm{full}}\) is the finite sum \[\Sigma_{\mathrm{full}} \;=\; \sum_{\mathbf{d}, \mathbf{e}}\ \sum_{\substack{u_1, \dotsc, u_k\\ u_i \mid (d_i, e_i)}}\ \sum_{\substack{s_{1,2}, \dotsc, s_{k, k-1}\\ s_{i, j} \mid (d_i, e_j)\ (i \ne j)}} \Bigl(\prod_{i=1}^k \varphi(u_i)\Bigr) \Bigl(\prod_{i \ne j} \mu(s_{i, j})\Bigr) \frac{\lambda(\mathbf{d})\, \lambda(\mathbf{e})}{\prod_{i=1}^k d_i e_i}.\]
Proof. Step 1: factor the indicator. The constraint is a conjunction over the \(k(k-1)\) ordered pairs \((i, j)\) with \(i \ne j\), so \[\mathbf{1}\bigl[(d_i, e_j) = 1\ \forall\, i \ne j\bigr] \;=\; \prod_{i \ne j} \mathbf{1}\bigl[(d_i, e_j) = 1\bigr].\]
Step 2: detect each factor by Möbius inversion. By Lemma 1 applied at \(n = (d_i, e_j)\), \[\mathbf{1}\bigl[(d_i, e_j) = 1\bigr] \;=\; \sum_{s_{i, j} \mid (d_i, e_j)} \mu(s_{i, j}),\] and \(s_{i, j} \mid (d_i, e_j)\) is the conjunction \(s_{i, j} \mid d_i\) and \(s_{i, j} \mid e_j\). Taking the product over the \(k(k-1)\) pairs turns Step 1 into a single sum over tuples \((s_{i, j})_{i \ne j}\) of positive integers weighted by \(\prod_{i \ne j} \mu(s_{i, j})\).
Step 3: insert and interchange. Insert the expansion of Step 2 into the left-hand side. All sums are finite: \(\mathbf{d}, \mathbf{e}\) range over the finite support of \(\lambda\), the \(u_i\) over divisors of \((d_i, e_i)\) and the \(s_{i, j}\) over divisors of \((d_i, e_j)\). Interchanging the order of summation and merging the divisibility indicators into the summation ranges yields exactly \(\Sigma_{\mathrm{full}}\). ◻
Proof uses: lem_mobius_divisor_sum
Lemma 1 (Pointwise coprimalities of the \(s_{i, j}\)). Let \(\lambda\) be such that \(\lambda(\mathbf{t}) \ne 0\) implies that every \(t_i\) is positive and that \(t_1, \dotsc, t_k\) are pairwise coprime. Let \((\mathbf{d}, \mathbf{e}, \mathbf{u}, (s_{i, j})_{i \ne j})\) be a configuration occurring in \(\Sigma_{\mathrm{full}}\), so that \(u_i \mid d_i\), \(u_i \mid e_i\) and \(s_{i, j} \mid d_i\), \(s_{i, j} \mid e_j\) for \(i \ne j\), and suppose its summand is nonzero, i.e. \(\lambda(\mathbf{d})\lambda(\mathbf{e}) \ne 0\). Then the four restrictions \((\ast_1)\)–\((\ast_4)\) of Notation 1 hold.
Proof. By hypothesis the coordinates of \(\mathbf{d}\) are pairwise coprime and so are those of \(\mathbf{e}\): \[(d_i, d_j) = 1, \qquad (e_i, e_j) = 1 \qquad (i \ne j). \tag{$\ast$}\] Each of the four cases assumes a prime \(p\) divides \(s_{i, j}\) and one further listed factor, and contradicts \((\ast)\).
\((\ast_1)\). If \(p \mid s_{i, j}\) and \(p \mid u_i\), then \(p \mid s_{i, j} \mid e_j\) and \(p \mid u_i \mid e_i\), so \(p \mid (e_i, e_j)\) with \(i \ne j\): contradiction.
\((\ast_2)\). If \(p \mid s_{i, j}\) and \(p \mid u_j\), then \(p \mid s_{i, j} \mid d_i\) and \(p \mid u_j \mid d_j\), so \(p \mid (d_i, d_j)\): contradiction.
\((\ast_3)\). If \(p \mid s_{i, j}\) and \(p \mid s_{i, a}\) with \(a \notin \{i, j\}\), then \(p \mid e_j\) and \(p \mid e_a\) with \(j \ne a\): contradiction.
\((\ast_4)\). If \(p \mid s_{i, j}\) and \(p \mid s_{b, j}\) with \(b \notin \{i, j\}\), then \(p \mid d_i\) and \(p \mid d_b\) with \(i \ne b\): contradiction. ◻
Lemma 1 (The restricted and unrestricted \(s\)-sums agree). Under the hypothesis of Lemma 1 on \(\lambda\), let \(\Sigma_{\mathrm{restr}}\) denote the sum obtained from \(\Sigma_{\mathrm{full}}\) by replacing the unrestricted \(s\)-summation by the restricted summation \(\sideset{}{^*}\sum\) of Notation 1. Then \[\Sigma_{\mathrm{restr}} \;=\; \Sigma_{\mathrm{full}}.\]
Proof. The restricted sum is obtained from the unrestricted one by deleting the configurations that violate at least one of \((\ast_1)\)–\((\ast_4)\). By Lemma 1, every configuration whose summand is nonzero satisfies all four restrictions; contrapositively, every deleted configuration has summand \(0\). Deleting terms that vanish does not change the value of a sum, so the two sums are equal. ◻
The substitution \(\lambda \rightsquigarrow y\)
Lemma 1 (Substituting \(y\) into \(S_1(\lambda)\)). Let \(k \ge 2\), let \(\mathcal{H}\) be admissible, let \(\vartheta \in (0, 1)\) and \(\delta \in (0, \vartheta/2)\). Then, for all sufficiently large \(N\), for every \(\lambda\) with \(R\)-permissible support and every valid \(v_0\), \[S_1(\lambda) \;=\; \frac{N}{W}\,\Sigma_y(\lambda) \;+\; O_k\Bigl(\frac{y_{\max}^2\, \varphi(W)^k\, N\, (\log R)^k}{W^{k+1}\, D_0}\Bigr),\] where, with \(\mathbf{a}, \mathbf{b}\) as in Definitions 1 and 1, \(\Sigma_y(\lambda)\) is the finite sum \[\Sigma_y(\lambda) \;=\; \sum_{\mathbf{u}} \Bigl(\prod_{i=1}^k \frac{\mu(u_i)^2}{\varphi(u_i)}\Bigr) \sideset{}{^*}\sum_{s_{1,2}, \dotsc, s_{k, k-1}} \Bigl(\prod_{i \ne j} \frac{\mu(s_{i, j})}{\varphi(s_{i, j})^2}\Bigr)\, y(a_1, \dotsc, a_k)\, y(b_1, \dotsc, b_k).\]
Uses: def_S1, def_admissible, def_adm_support, def_R, def_D_0, def_W_trick, def_y_from_lambda, def_y_max, def_bold_a, def_bold_b, lem_exists_v0, not_starsum, not_bold_tuple, not_standing
Proof. Step 1: the starting point. By Lemma 1, \(S_1(\lambda)\) equals \((N/W)\) times the primed sum, up to the stated error. Applying, in turn, Lemma 1 (legitimate: \(R\)-permissibility of the support gives positivity, squarefreeness of \(\prod_i t_i\) and coprimality to \(W\)), Lemma 1 (legitimate: squarefreeness of \(\prod_i t_i\) gives squarefreeness and pairwise coprimality of the coordinates), Lemma 1 and Lemma 1, we obtain \[S_1(\lambda) \;=\; \frac{N}{W} \sum_{\mathbf{u}} \Bigl(\prod_{i=1}^k \varphi(u_i)\Bigr) \sideset{}{^*}\sum_{s_{1,2}, \dotsc, s_{k, k-1}} \Bigl(\prod_{i \ne j} \mu(s_{i, j})\Bigr) \sum_{\substack{\mathbf{d}, \mathbf{e}\\ u_i \mid d_i, e_i\\ s_{i, j} \mid d_i,\, e_j}} \frac{\lambda(\mathbf{d})\, \lambda(\mathbf{e})}{\prod_{i=1}^k d_i e_i} \;+\; E,\] with \(E\) the error term of Lemma 1.
Step 2: collapse the \(\mathbf{d}\)-sum. Fix \(\mathbf{u}\) and \((s_{i, j})\) satisfying the restrictions of Notation 1, which by Lemma 1 is no loss. Then, for each \(i\), the integers \(u_i\) and \(\{s_{i, j}\}_{j \ne i}\) are pairwise coprime, so the conjunction of \(u_i \mid d_i\) and \(s_{i, j} \mid d_i\) for all \(j \ne i\) is equivalent to the single divisibility \(a_i \mid d_i\) (Definition 1). Hence the inner \(\mathbf{d}\)-sum is over \(a_i \mid d_i\) for all \(i\), and by Definition 1 and the Möbius inversion of Lemma 1, \[\sum_{\substack{\mathbf{d}\\ a_i \mid d_i\ \forall i}} \frac{\lambda(\mathbf{d})}{\prod_i d_i} \;=\; \frac{y(a_1, \dotsc, a_k)}{\prod_{i=1}^k \mu(a_i)\, \varphi(a_i)}\] whenever \(\prod_i a_i\) is squarefree, the summand vanishing otherwise. The same argument with \(b_i\) in place of \(a_i\) (using coprimality of \(u_i\) and \(\{s_{j, i}\}_{j \ne i}\), and Definition 1) collapses the \(\mathbf{e}\)-sum.
Step 3: recombine the arithmetic factors. On the support, \(a_i = u_i \prod_{j \ne i} s_{i, j}\) and \(b_i = u_i \prod_{j \ne i} s_{j, i}\) are squarefree with pairwise coprime factors, so \(\mu\) and \(\varphi\) are multiplicative across them. Collecting the three sources of \(\mu(s_{i, j})\) — one from \(\mu(a_i)\), one from \(\mu(b_j)\), and the Möbius weight inserted in Lemma 1 — and the factors \(\varphi(u_i)^2\) against \(\prod_i \varphi(u_i)\), we obtain, for each \(i\), a factor \(\mu(u_i)^2/\varphi(u_i)\), and for each ordered pair \(i \ne j\) a factor \(\mu(s_{i, j})/\varphi(s_{i, j})^2\). This is exactly the summand of \(\Sigma_y(\lambda)\), and Step 1’s error term is unchanged. ◻
Lemma 1 (Contributing \(s_{i, j}\) are coprime to \(W\)). Let \(\lambda\) have \(R\)-permissible support. If a configuration \((\mathbf{u}, (s_{i, j}))\) contributes to \(\Sigma_y(\lambda)\), that is if \(y(a_1, \dotsc, a_k) \ne 0\) and \(y(b_1, \dotsc, b_k) \ne 0\), then \((s_{i, j}, W) = 1\) for every \(i \ne j\).
Proof. Fix \(i \ne j\). By Definition 1, \(y(b_1, \dotsc, b_k) \ne 0\) forces the inner sum defining it to be nonzero, so there is a tuple \(\mathbf{d}\) with \(\lambda(\mathbf{d}) \ne 0\) and \(b_\ell \mid d_\ell\) for every \(\ell\). By Definition 1, \(s_{i, j} \mid b_j \mid d_j\). By \(R\)-permissibility of the support of \(\lambda\) we have \((\prod_\ell d_\ell, W) = 1\), hence \((d_j, W) = 1\) and therefore \((s_{i, j}, W) = 1\). ◻
Lemma 1 (Dichotomy for the contributing \(s_{i, j}\)). Let \(\lambda\) have \(R\)-permissible support and let \(W = \prod_{p \le D_0} p\). If a configuration contributes to \(\Sigma_y(\lambda)\), then for every \(i \ne j\) either \(s_{i, j} = 1\), or every prime factor of \(s_{i, j}\) exceeds \(D_0\); in particular \(s_{i, j} = 1\) or \(s_{i, j} > D_0\).
Proof. If \(s_{i, j} = 1\) there is nothing to prove, so assume \(s_{i, j} > 1\) and let \(q\) be a prime factor of \(s_{i, j}\). By Lemma 1, \((s_{i, j}, W) = 1\), so \(q \nmid W\). By Definition 1, \(W\) is the product of all primes \(p \le D_0\), so a prime divides \(W\) if and only if it is at most \(D_0\). Hence \(q > D_0\). Since \(s_{i, j}\) is then divisible by a prime exceeding \(D_0\), we get \(s_{i, j} > D_0\). ◻
Lemma 1 (The contribution of \(s_{i, j} > D_0\)). Let \(k \ge 2\), \(\vartheta \in (0, 1)\), \(\delta \in (0, \vartheta/2)\). For all sufficiently large \(N\) and every \(\lambda\) with \(R\)-permissible support, the part of \((N/W)\Sigma_y(\lambda)\) supported on configurations with \(s_{i, j} > D_0\) for at least one pair \(i \ne j\) is \[O_k\Bigl(\frac{y_{\max}^2\, \varphi(W)^k\, N\, (\log R)^k}{W^{k+1}\, D_0}\Bigr).\]
Proof. There are \(k(k-1)\) ordered pairs \((i, j)\) with \(i \ne j\), so by the union bound it suffices to bound, for one fixed pair, the contribution of the configurations with \(s_{i, j} > D_0\), and then multiply by \(k(k-1)\).
Bound \(|y(\mathbf{a})\, y(\mathbf{b})| \le y_{\max}^2\) and take absolute values in the summand. Every variable is constrained to be squarefree and coprime to \(W\) (Lemma 1 and \(R\)-permissibility), and each of \(\mathbf{a}, \mathbf{b}\) has product at most \(R\) on the support of \(y\), so each \(u_\ell\) and each \(s_{a, b}\) is at most \(R\). Therefore the contribution of the fixed pair is at most \[y_{\max}^2 \Bigl(\sum_{\substack{u \le R\\ (u, W) = 1}} \frac{\mu(u)^2}{\varphi(u)}\Bigr)^{\!k} \Bigl(\sum_{\substack{s \le R\\ (s, W) = 1}} \frac{\mu(s)^2}{\varphi(s)^2}\Bigr)^{\!k(k-1) - 1} \sum_{\substack{s > D_0\\ (s, W) = 1}} \frac{\mu(s)^2}{\varphi(s)^2}.\] By Lemma 1 the first factor is \(\ll \bigl(\varphi(W)(\log R)/W\bigr)^k\); the middle factor is \(O_k(1)\) because \(\sum_{s \ge 1} \mu(s)^2/\varphi(s)^2\) converges; and the last factor is \(\ll 1/D_0\), since \(\varphi(s) \gg s/\log\log(3s)\) makes the tail \(\sum_{s > D_0} \mu(s)^2/\varphi(s)^2\) comparable to \(\sum_{s > D_0} s^{-2 + o(1)} \ll 1/D_0\). Multiplying by \(N/W\) and by \(k(k-1)\) gives the assertion. ◻
Proof uses: lem_S1_sij_coprime_to_W, lem_mertens_W
Lemma 1 (The diagonal term \(s_{i, j} \equiv 1\)). For every \(\lambda : \mathbb{N}^k \to \mathbb{R}\) of finite support, \[\sum_{\mathbf{u}} \Bigl(\prod_{i=1}^k \frac{\mu(u_i)^2}{\varphi(u_i)}\Bigr)\, y(u_1, \dotsc, u_k)^2 \;=\; \sum_{\mathbf{u}} \frac{y(u_1, \dotsc, u_k)^2}{\prod_{i=1}^k \varphi(u_i)},\] the sums being over \(k\)-tuples of positive integers. (The left-hand side is the summand of \(\Sigma_y(\lambda)\) at \(s_{i, j} = 1\) for all \(i \ne j\), since then \(\mathbf{a} = \mathbf{b} = \mathbf{u}\), \(\mu(1) = \varphi(1) = 1\), and the restrictions of Notation 1 are vacuous.)
Uses: lem_S1_substitute_y, def_y_from_lambda, def_mobius, def_totient, def_bold_a, def_bold_b, not_starsum
Proof. It suffices to compare the two summands at a fixed tuple \(\mathbf{u}\). If \(\prod_i u_i\) is squarefree then each \(\mu(u_i)^2 = 1\) and the two summands coincide. If \(\prod_i u_i\) is not squarefree, then by Definition 1 the value \(y(u_1, \dotsc, u_k)\) vanishes — the \(y\)-variables are supported on tuples with squarefree product — so both summands are \(0\). In either case the summands agree, and hence so do the sums. ◻
Lemma 1 (\(S_1(\lambda)\) in terms of \(y\)). Let \(k \ge 2\), let \(\mathcal{H}\) be admissible, let \(\vartheta \in (0, 1)\) and \(\delta \in (0, \vartheta/2)\). Then, for all sufficiently large \(N\), for every \(\lambda\) with \(R\)-permissible support and every valid \(v_0\), \[S_1(\lambda) \;=\; \frac{N}{W} \sum_{\mathbf{u}} \frac{y(u_1, \dotsc, u_k)^2}{\prod_{i=1}^k \varphi(u_i)} \;+\; O_k\Bigl(\frac{y_{\max}^2\, \varphi(W)^k\, N\, (\log R)^k}{W^{k+1}\, D_0}\Bigr).\]
Uses: def_S1, def_admissible, def_adm_support, def_R, def_D_0, def_W_trick, def_y_from_lambda, def_y_max, lem_exists_v0, not_bold_tuple, not_standing
Proof. By Lemma 1, \(S_1(\lambda) = (N/W)\,\Sigma_y(\lambda) + E_1\) with \(E_1 \ll_k y_{\max}^2 \varphi(W)^k N (\log R)^k/(W^{k+1} D_0)\).
Split \(\Sigma_y(\lambda)\) according to whether all the \(s_{i, j}\) with \(i \ne j\) equal \(1\). By Lemma 1, on every contributing configuration each \(s_{i, j}\) is either \(1\) or exceeds \(D_0\); hence the configurations that are not diagonal are exactly those with \(s_{i, j} > D_0\) for at least one pair, and their total contribution to \((N/W)\Sigma_y(\lambda)\) is \(O_k\bigl(y_{\max}^2 \varphi(W)^k N (\log R)^k/(W^{k+1} D_0)\bigr)\) by Lemma 1.
The diagonal part is \((N/W)\) times the left-hand side of Lemma 1, which that lemma identifies with \((N/W)\sum_{\mathbf{u}} y(\mathbf{u})^2/\prod_i \varphi(u_i)\). Adding the two error contributions gives the assertion. ◻
Proof uses: lem_S1_substitute_y, lem_S1_sij_dichotomy, lem_S1_sij_D0, lem_S1_sij_one
The smooth case
Lemma 1 (Substituting the smooth \(y\)). Let \(k \ge 2\), let \(\mathcal{H}\) be admissible, let \(\vartheta \in (0, 1)\) and \(\delta \in (0, \vartheta/2)\), and let \(F : \mathbb{R}^k \to \mathbb{R}\) be smooth with support contained in \(\mathcal{R}_k\) (Definition 1). Let \(\lambda_0\) be the \(F\)-derived sieve weight of Definition 1. Then there is a constant \(C > 0\), depending only on \(k\), \(\vartheta\) and \(\delta\), such that for all sufficiently large \(N\) and every valid \(v_0\), \[\Bigl|S_1(\lambda_0) \;-\; \frac{N}{W} \sum_{\substack{u_1, \dotsc, u_k\\ \prod_i u_i \ \mathrm{squarefree}\\ (\prod_i u_i, W) = 1}} \frac{1}{\prod_{i=1}^k \varphi(u_i)}\, F\Bigl(\frac{\log u_1}{\log R}, \dotsc, \frac{\log u_k}{\log R}\Bigr)^{\!2}\Bigr| \;\le\; C\,\frac{F_{\max}^2\, \varphi(W)^k\, N\, (\log R)^k}{W^{k+1}\, D_0}.\]
Uses: lem_S1_from_y, lem_smooth_y, def_lambda_from_F, def_simplex, def_F_smooth_max, def_y_from_lambda, def_y_max, def_R, def_D_0, def_W_trick, not_standing
Proof. Apply Lemma 1 to \(\lambda = \lambda_0\), which has \(R\)-permissible support by Definition 1. By Lemma 1, the associated \(y\)-variables are given in closed form: \[y(u_1, \dotsc, u_k) \;=\; \begin{cases} F\bigl(\tfrac{\log u_1}{\log R}, \dotsc, \tfrac{\log u_k}{\log R}\bigr) & \text{if } \prod_i u_i \text{ is squarefree and } (\prod_i u_i, W) = 1,\\ 0 & \text{otherwise.} \end{cases}\] Substituting this into the main term of Lemma 1 restricts the \(\mathbf{u}\)-sum to tuples with \(\prod_i u_i\) squarefree and coprime to \(W\) and replaces \(y(\mathbf{u})^2\) by \(F(\log u_1/\log R, \dotsc)^2\), which is the displayed main term.
For the error term, the same closed form gives \(y_{\max} = \sup_{\mathbf{u}} |y(\mathbf{u})| \le \sup_{[0,1]^k} |F| \le F_{\max}\): indeed the support condition on \(F\) forces \(F\) to vanish unless \(\log u_i/\log R \in [0, 1]\) for every \(i\), and \(F_{\max} \ge \sup_{[0,1]^k}|F|\) by Definition 1. Replacing \(y_{\max}\) by \(F_{\max}\) in the error term of Lemma 1 gives the stated bound. ◻
Proof uses: lem_S1_from_y, lem_smooth_y
Lemma 1 (Dropping the constraint \((u_i, u_j) = 1\)). Let \(\vartheta \in (0, 1)\), \(\delta \in (0, \vartheta/2)\). There is a constant \(C > 0\) depending only on \(k\), \(\vartheta\), \(\delta\) such that for all sufficiently large \(N\) and every smooth \(F : \mathbb{R}^k \to \mathbb{R}\) supported in \(\mathcal{R}_k\), \[\Bigl| \sum_{\substack{\mathbf{u}\\ \prod_i u_i\ \mathrm{squarefree}\\ (\prod_i u_i, W) = 1}} \frac{F\bigl(\tfrac{\log u_1}{\log R}, \dotsc\bigr)^2}{\prod_i \varphi(u_i)} \;-\; \sum_{\substack{\mathbf{u}\\ (u_i, W) = 1\ \forall i}} \Bigl(\prod_{i=1}^k \frac{\mu(u_i)^2}{\varphi(u_i)}\Bigr) F\Bigl(\tfrac{\log u_1}{\log R}, \dotsc\Bigr)^{\!2} \Bigr| \;\le\; C\,\frac{F_{\max}^2\, \varphi(W)^k\, (\log R)^k}{W^k\, D_0}.\]
Uses: lem_S1_substitute_smooth_y, def_simplex, def_F_smooth_max, def_R, def_D_0, def_W_trick, def_mobius, def_totient
Proof. Step 1: identify the difference. A tuple \(\mathbf{u}\) has \(\prod_i u_i\) squarefree and coprime to \(W\) if and only if every \(u_i\) is squarefree and coprime to \(W\) and the \(u_i\) are pairwise coprime. On such tuples \(\mu(u_i)^2 = 1\) for every \(i\), so the two summands coincide. The second sum runs over the larger index set \(\{(u_i, W) = 1\ \forall i\}\), on which its summand is nonnegative and vanishes unless each \(u_i\) is squarefree. Hence the difference is, in absolute value, exactly \[\sum_{\substack{\mathbf{u}:\ (u_i, W) = 1\ \forall i\\ \exists\, i \ne j:\ (u_i, u_j) > 1}} \Bigl(\prod_{\ell=1}^k \frac{\mu(u_\ell)^2}{\varphi(u_\ell)}\Bigr) F\Bigl(\tfrac{\log u_1}{\log R}, \dotsc\Bigr)^{\!2}.\]
Step 2: a shared prime is large. Fix \(i \ne j\) and suppose a prime \(p\) divides both \(u_i\) and \(u_j\). Since \((u_i, W) = 1\) and \(W = \prod_{q \le D_0} q\), we have \(p > D_0\).
Step 3: bound one pair. Bound \(F^2 \le F_{\max}^2\); the support of \(F\) inside \(\mathcal{R}_k\) forces \(u_\ell \le R\) for every \(\ell\) in each nonzero term. For fixed \(i \ne j\) and a fixed prime \(p > D_0\), summing over \(u_i, u_j\) divisible by \(p\) and over the remaining coordinates freely, \[\sum_{\substack{\mathbf{u} \le R,\ (u_\ell, W) = 1\\ p \mid u_i,\ p \mid u_j}} \prod_\ell \frac{\mu(u_\ell)^2}{\varphi(u_\ell)} \;\le\; \frac{1}{(p-1)^2}\, \Bigl(\sum_{\substack{u \le R\\ (u, W) = 1}} \frac{\mu(u)^2}{\varphi(u)}\Bigr)^{\!k},\] because for squarefree \(u_i\) with \(p \mid u_i\) one has \(\mu(u_i)^2/\varphi(u_i) = (p-1)^{-1}\mu(u_i/p)^2/\varphi(u_i/p)\).
Step 4: sum over \(p\) and over pairs. By Lemma 1, \(\sum_{u \le R, (u, W) = 1} \mu(u)^2/\varphi(u) \ll \varphi(W)(\log R)/W\). Moreover \(\sum_{p > D_0} (p-1)^{-2} \ll 1/D_0\). Summing Step 3 over the primes \(p > D_0\) and over the \(k(k-1)\) ordered pairs \((i, j)\) bounds the quantity of Step 1 by \[\ll_k\ F_{\max}^2 \cdot \frac{1}{D_0} \cdot \frac{\varphi(W)^k (\log R)^k}{W^k},\] which is the assertion. ◻
Proof uses: lem_mertens_W
Lemma 1 (Applying partial summation once per coordinate). Let \(\vartheta \in (0, 1)\), \(\delta \in (0, \vartheta/2)\). There is a constant \(C > 0\) depending only on \(k\), \(\vartheta\), \(\delta\) such that for all sufficiently large \(N\) and every smooth \(F : \mathbb{R}^k \to \mathbb{R}\) supported in \(\mathcal{R}_k\), \[\begin{aligned} &\Bigl| \sum_{\substack{u_1, \dotsc, u_k\\ (u_i, W) = 1\ \forall i}} \Bigl(\prod_{i=1}^k \frac{\mu(u_i)^2}{\varphi(u_i)}\Bigr) F\Bigl(\frac{\log u_1}{\log R}, \dotsc, \frac{\log u_k}{\log R}\Bigr)^{\!2} \;-\; \frac{\varphi(W)^k (\log R)^k}{W^k}\, I_k(F) \Bigr|\\ &\;\le\; C\,\frac{F_{\max}^2\, \varphi(W)^k\, (\log D_0)\, (\log R)^{k-1}}{W^k}. \end{aligned}\]
Uses: lem_S1_drop_uij, lem_partial_sum, def_simplex, def_I_k, def_F_smooth_max, def_R, def_D_0, def_W_trick, not_singular_series
Proof. We verify the hypotheses of Lemma 1 for the density \(\gamma_W\) and then iterate it once per coordinate.
Step 1: the sieve data. Let \(\gamma_W\) be the multiplicative density with \(\gamma_W(p) = \mathbf{1}[p \nmid W]\). Then \(\gamma_W(p)/p \le 1/2\) for all \(p\), so the growth hypothesis holds with \(A_1 = 1/2\). For the logarithmic hypothesis, Mertens’ first theorem gives \(\sum_{w \le p \le z} (\log p)/p = \log(z/w) + O(1)\), while the primes dividing \(W\) contribute at most \(\sum_{p \le D_0} (\log p)/p = \log D_0 + O(1)\); subtracting, the hypothesis holds with \(A_2 = O(1)\) and \(L \ll \log D_0\). Finally the structural hypothesis holds with \(V = W\) (squarefree) and \(A_3 = 0\), since \(\gamma_W(p) = 0\) for \(p \mid W\) and \(\gamma_W(p) = 1\) otherwise; the associated residual error of Lemma 1 is \(O(\tau(W)\log(2WR)/\log R)\), negligible because \(W \le (\log\log N)^2\) while \(\log R\) is a fixed positive multiple of \(\log N\).
Step 2: the singular product and the multiplicative weight. By Lemma 1, \(\mathfrak{S}(\gamma_W) = \varphi(W)/W\). The associated totally multiplicative function is \(g_*(p) = \gamma_W(p)/(p - \gamma_W(p))\), i.e. \(g_*(p) = 0\) for \(p \mid W\) and \(g_*(p) = 1/(p-1)\) for \(p \nmid W\). Consequently, for squarefree \(u\) coprime to \(W\), \(g_*(u) = 1/\varphi(u)\), and \(\mu(u)^2 g_*(u) = \mu(u)^2/\varphi(u)\) is supported exactly on the squarefree \(u\) coprime to \(W\) — the vanishing off the integers coprime to the modulus being Lemma 1 at \(V = W\) — which is precisely the summand appearing in the statement.
Step 3: one coordinate at a time. Take \(z := R\) throughout; the support of \(F\) inside \(\mathcal{R}_k\) makes the summand vanish unless \(\log u_i/\log R \le 1\), i.e. \(u_i \le R\). Fix values \(x_\ell := \log u_\ell/\log R\) for \(\ell > 1\) and apply Lemma 1 to the smooth profile \(G_1(t) := F(t, x_2, \dotsc, x_k)^2\) on \([0, 1]\), whose \(C^1\)-norm satisfies \(G_{\max} \ll F_{\max}^2\) by Definition 1. This yields \[\begin{aligned} \sum_{\substack{u_1 \le R\\ (u_1, W) = 1}} \frac{\mu(u_1)^2}{\varphi(u_1)} F\Bigl(\frac{\log u_1}{\log R}, x_2, \dotsc, x_k\Bigr)^{\!2} &\;=\; \frac{\varphi(W)}{W}(\log R)\int_0^1 F(t_1, x_2, \dotsc, x_k)^2\, dt_1\\ &\;+\; O\Bigl(\frac{\varphi(W)}{W}(\log D_0)\, F_{\max}^2\Bigr). \end{aligned}\] Iterate over the coordinates \(i = 2, \dotsc, k\), at stage \(i\) applying Lemma 1 to the profile \(G_i(t) := \int_{[0,1]^{i-1}} F(t_1, \dotsc, t_{i-1}, t, x_{i+1}, \dotsc, x_k)^2\, dt_1 \dotsm dt_{i-1}\), whose \(C^1\)-norm is again \(\ll F_{\max}^2\).
Step 4: aggregate. After \(k\) iterations the main term is \[\Bigl(\frac{\varphi(W)\log R}{W}\Bigr)^{\!k} \int_{[0,1]^k} F(t_1, \dotsc, t_k)^2\, dt_1 \dotsm dt_k \;=\; \frac{\varphi(W)^k (\log R)^k}{W^k}\, I_k(F),\] using that \(F\) is supported in \(\mathcal{R}_k \subseteq [0, 1]^k\) so that the integral over the cube is \(I_k(F)\) (Definition 1). The total error is the sum over \(i = 1, \dotsc, k\) of the terms in which the error appears in the \(i\)-th coordinate and the main term in the other \(k - 1\); each such term has size \[\Bigl(\frac{\varphi(W)\log R}{W}\Bigr)^{\!k-1} \cdot \frac{\varphi(W)}{W}(\log D_0)\, F_{\max}^2 \;=\; \frac{F_{\max}^2\, \varphi(W)^k (\log D_0)(\log R)^{k-1}}{W^k},\] and there are \(k\) of them. This is the stated bound. ◻
Proof uses: lem_partial_sum, slem_gg_singular_series, lem_h_support, lem_W_size
Lemma 1 (Asymptotic for \(S_1(\lambda_0)\) in the smooth case). Let \(k \ge 2\), let \(\mathcal{H}\) be admissible, let \(\vartheta \in (0, 1)\) and \(\delta \in (0, \vartheta/2)\), and let \(F : \mathbb{R}^k \to \mathbb{R}\) be smooth with support contained in \(\mathcal{R}_k\). Let \(\lambda_0\) be the \(F\)-derived sieve weight of Definition 1. Then there is a constant \(C > 0\), depending only on \(k\), \(\vartheta\) and \(\delta\), such that for all sufficiently large \(N\) and every valid \(v_0\), \[\Bigl|S_1(\lambda_0) \;-\; \frac{\varphi(W)^k\, N\, (\log R)^k}{W^{k+1}}\, I_k(F)\Bigr| \;\le\; C\,\frac{F_{\max}^2\, \varphi(W)^k\, N\, (\log R)^k}{W^{k+1}\, D_0}.\]
Uses: def_S1, def_admissible, def_lambda_from_F, def_simplex, def_I_k, def_F_smooth_max, def_R, def_D_0, def_W_trick, lem_exists_v0, not_standing
Proof. Write \(A := \sum_{\mathbf{u}} \mathbf{1}[\prod_i u_i \text{ squarefree}, (\prod_i u_i, W) = 1]\, F(\log u_1/\log R, \dotsc)^2 / \prod_i \varphi(u_i)\) for the sum appearing in Lemma 1, \(B\) for the sum appearing on the right of Lemma 1, and \(M := \varphi(W)^k (\log R)^k I_k(F)/W^k\). The three preceding lemmas give, for \(N\) large, \[\Bigl|S_1(\lambda_0) - \frac{N}{W} A\Bigr| \le C_1 \frac{F_{\max}^2 \varphi(W)^k N (\log R)^k}{W^{k+1} D_0}, \qquad |A - B| \le C_2 \frac{F_{\max}^2 \varphi(W)^k (\log R)^k}{W^k D_0},\] \[|B - M| \le C_3 \frac{F_{\max}^2 \varphi(W)^k (\log D_0)(\log R)^{k-1}}{W^k}.\] Since \(N/W \ge 0\), the triangle inequality gives \[\Bigl|S_1(\lambda_0) - \frac{N}{W} M\Bigr| \;\le\; C_1 \frac{F_{\max}^2 \varphi(W)^k N (\log R)^k}{W^{k+1} D_0} \;+\; \frac{N}{W}\bigl(|A - B| + |B - M|\bigr),\] and \((N/W)M\) is exactly \(\varphi(W)^k N (\log R)^k I_k(F)/W^{k+1}\), the asserted main term.
It remains to see that the third error is of the same shape as the first two. Set \(K := F_{\max}^2 \varphi(W)^k N (\log R)^{k-1}/W^{k+1} \ge 0\); then the first two errors are \(C_1 K (\log R)/D_0\) and \(C_2 K (\log R)/D_0\), while the third is \(C_3 K \log D_0\). Since \(D_0 = \log\log\log N\) and \(\log R = (\vartheta/2 - \delta)\log N\), we have \(D_0 \log D_0 \le \log R\) for all large \(N\), that is \(\log D_0 \le (\log R)/D_0\). Hence the third error is at most \(C_3 K (\log R)/D_0\), and the total is at most \((C_1 + C_2 + C_3) K (\log R)/D_0\), which is the stated bound with \(C := C_1 + C_2 + C_3\). ◻
Sieve manipulations: the reduction of \(S_2^{(m)}\)
For each \(m \in \{1, \dotsc, k\}\) we carry out the passage from the sieve coefficients \(\lambda\) to the smooth \(y\)-variables in the second moment \(S_2^{(m)}(\lambda)\), ending in an asymptotic expressed through the functional \(J_k^{(m)}\). The argument parallels the reduction of \(S_1(\lambda)\) — expand the square, evaluate the inner sum by the Chinese remainder theorem, decouple the least common multiples, invert the residual coprimality conditions by Möbius, and substitute the transformed variables — but two features are new. First, the inner sum now carries the prime indicator \(\chi_{\mathbb{P}}(n + h_m)\), so the Chinese remainder step produces a main term \(X_N/\varphi(q)\) and an error \(E(N, q)\) rather than an exact count; controlling the aggregate of those errors over the whole support of \(\lambda\) requires the hypothesis that the primes have a level of distribution \(\vartheta\), and is the only place where Bombieri–Vinogradov (Theorem 1) is used. Second, the transformed variables are the \(y^{(m)}\) of Definition 1, whose normalisation is \(\prod_i \varphi(d_i)\) rather than \(\prod_i d_i\); this is the source of the multiplicative function \(g\) of Definition 1 throughout.
Throughout the section \(k \ge 2\) is fixed, \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) is admissible, \(m \in \{1, \dotsc, k\}\) is a distinguished index, \(D_0 = \log\log\log N\), \(W = \prod_{p \le D_0} p\), \(v_0\) is a residue compatible with \(\mathcal{H}\) modulo \(W\), \(\vartheta \in (0, 1/2)\) is a level of distribution for the primes, \(\delta \in (0, \vartheta/2)\) and \(R = N^{\vartheta/2 - \delta}\). The sieve coefficients \(\lambda\) are always assumed to have permissible support (Definition 1) at the parameters \(R\) and \(W\). All implied constants depend only on \(k\) and \(\mathcal{H}\), never on \(N\), on \(\lambda\), on the profile \(F\), or on \(m\), unless explicitly stated otherwise.
Expansion and the Chinese remainder step
Lemma 1 (Expansion of \(S_2^{(m)}(\lambda)\)). For every finitely supported \(\lambda\), \[S_2^{(m)}(\lambda) \;=\; \sum_{\mathbf{d}, \mathbf{e}} \lambda(\mathbf{d})\, \lambda(\mathbf{e}) \sum_{\substack{N < n \le 2N\\ n \equiv v_0 \!\!\pmod W\\ [d_i, e_i] \mid n + h_i\ \forall i}} \chi_{\mathbb{P}}(n + h_m).\]
Uses: def_S2m, def_weight, def_W_trick
Proof. This is the same three-step rearrangement that produced the expansion of \(S_1(\lambda)\) (Lemma 1), with the extra factor \(\chi_{\mathbb{P}}(n + h_m)\) carried inertly through every step; only the coordinatewise divisibility identity is used, so the argument applies verbatim to any weight attached to \(n\).
Step 1: expand the square. By Definition 1, \[w_n \;=\; \Bigl(\sum_{d_i \mid n + h_i\, \forall i} \lambda(\mathbf{d})\Bigr)^{\!2} \;=\; \sum_{d_i \mid n + h_i\, \forall i}\ \sum_{e_i \mid n + h_i\, \forall i} \lambda(\mathbf{d})\,\lambda(\mathbf{e}),\] the two inner sums being finite because \(\lambda\) has finite support.
Step 2: substitute into \(S_2^{(m)}\). By Definition 1, \(S_2^{(m)}(\lambda)\) is the sum of \(\chi_{\mathbb{P}}(n + h_m)\, w_n\) over \(N < n \le 2N\) with \(n \equiv v_0 \pmod W\); inserting Step 1 gives a triple sum.
Step 3: interchange and merge the divisibilities. All three sums are finite, so the \(n\)-summation may be moved inside. For each fixed \(i\) the pair of conditions \(d_i \mid n + h_i\) and \(e_i \mid n + h_i\) holds if and only if \([d_i, e_i] \mid n + h_i\). Merging them coordinatewise yields the displayed identity. ◻
Proof uses: lem_S1_expansion
Lemma 1 (Primality forces \(d_m = e_m = 1\)). Let \(x, c\) be positive integers with \(2 \le c < x\) and \(c \mid x\). Then \(x\) is not prime. Consequently, for all \(N\) large enough that \(R < N\), the inner sum of Lemma 1 vanishes unless \(d_m = e_m = 1\).
Uses: lem_S2m_expansion, def_R, def_adm_support
Proof. The first assertion is immediate: \(c\) is a divisor of \(x\) with \(c \ne 1\) and \(c \ne x\), so \(x\) has a non-trivial divisor and is composite.
For the consequence, suppose \([d_m, e_m] > 1\) for a pair \((\mathbf{d}, \mathbf{e})\) in the support of \(\lambda\), and let \(n\) contribute to the inner sum, so that \([d_m, e_m] \mid n + h_m\). By Definition 1 the support of \(\lambda\) is contained in the tuples with \(\prod_i d_i \le R\), whence \([d_m, e_m] \le d_m e_m \le R^2\); more crudely, \([d_m, e_m] \le R < N < n + h_m\) once \(N\) exceeds a threshold depending only on \(\vartheta\) and \(\delta\). Applying the first assertion with \(c = [d_m, e_m]\) and \(x = n + h_m\) shows \(n + h_m\) is composite, so \(\chi_{\mathbb{P}}(n + h_m) = 0\). Hence the entire inner sum vanishes. ◻
Lemma 1 (Chinese remainder evaluation of the inner sum). There is a constant \(C > 0\), depending only on \(k\), \(\mathcal{H}\) and \(m\), with the following property. Let \(W \ge 1\), let \(v_0\) satisfy \((v_0 + h_i, W) = 1\) for every \(i\), and let \(\mathbf{d}, \mathbf{e}\) be tuples of positive integers such that
\(d_m = e_m = 1\);
\(W\) is coprime to \([d_i, e_i]\) for every \(i\), and the \([d_i, e_i]\) are pairwise coprime;
\((h_m - h_i,\, [d_i, e_i]) = 1\) for every \(i \ne m\).
Put \(q = q(\mathbf{d}, \mathbf{e}) = W \prod_{i=1}^k [d_i, e_i]\). Then \[\Bigl|\ \sum_{\substack{N < n \le 2N\\ n \equiv v_0 \!\!\pmod W\\ [d_i, e_i] \mid n + h_i\ \forall i}} \chi_{\mathbb{P}}(n + h_m) \;-\; \frac{X_N}{\varphi(q)}\ \Bigr| \;\le\; C\, E(N, q).\]
Uses: def_q_lcm, not_XN, not_error, def_totient, def_W_trick
Proof. Hypothesis (ii) makes \(W, [d_1, e_1], \dotsc, [d_k, e_k]\) a pairwise coprime family, so the Chinese remainder theorem identifies the system of congruences \[n \equiv v_0 \!\!\pmod W, \qquad n \equiv -h_i \!\!\pmod{[d_i, e_i]}\quad (1 \le i \le k)\] with a single congruence \(n \equiv a \pmod q\), for a residue \(a\) determined by the data. The inner sum is therefore \(\sum_{N < n \le 2N,\ n \equiv a\ (q)} \chi_{\mathbb{P}}(n + h_m)\), that is, the number of primes in \((N + h_m, 2N + h_m]\) lying in the single residue class \(a + h_m \pmod q\), up to at most \(2 h_k\) boundary terms; the boundary discrepancy is absorbed into the constant \(C\), using \(E(N, q) \ge 1\).
It remains to see that the class \(a + h_m\) is coprime to \(q\), so that the class is admissible for the definition of \(E(N, q)\) (Notation 1). Let \(p \mid q\). If \(p \mid W\) then \(a + h_m \equiv v_0 + h_m \pmod p\), which is coprime to \(W\) by the choice of \(v_0\). If \(p \mid [d_i, e_i]\) for some \(i\), then by (ii) this \(i\) is unique; when \(i = m\) hypothesis (i) gives \([d_m, e_m] = 1\), so no such \(p\) exists; when \(i \ne m\) we have \(a + h_m \equiv h_m - h_i \pmod p\), which is nonzero modulo \(p\) by hypothesis (iii). Hence \((a + h_m, q) = 1\).
By the very definition of \(E(N, q)\) as the supremum, over classes coprime to \(q\), of the discrepancy between the prime count in that class and \(X_N/\varphi(q)\), the inner sum differs from \(X_N/\varphi(q)\) by at most \(E(N, q)\), and the claim follows with a constant absorbing the boundary terms. ◻
The aggregate error term and Bombieri–Vinogradov
Lemma 1 replaces the inner sum by a main term at the cost of one error \(E(N, q(\mathbf{d}, \mathbf{e}))\) per pair \((\mathbf{d}, \mathbf{e})\). The aggregate of these errors, weighted by \(|\lambda(\mathbf{d})\lambda(\mathbf{e})|\), is negligible. Three ingredients combine: the moduli \(q(\mathbf{d}, \mathbf{e})\) are squarefree and lie below \(R^2 W\), which is a power \(N^{\vartheta - 2\delta} \cdot N^{o(1)}\) strictly below \(N^{1/2}\); each modulus is attained by at most \(\tau_{3k}\) pairs; and a level of distribution \(\vartheta\) controls the sum of \(E(N, q)\) over all \(q \le N^{\vartheta}\). Only the last of these uses anything beyond elementary arithmetic, and it is exactly Theorem 1.
Sublemma 1 (Squarefreeness of the sieve modulus). Let \(W\) be squarefree and let \((\mathbf{d}, \mathbf{e})\) be a pair in the support of a weight \(\lambda\) with permissible support at \((R, W)\). Then \(q(\mathbf{d}, \mathbf{e})\) is squarefree.
Uses: def_adm_support, def_q_lcm, def_W_trick
Proof. By Definition 1 the products \(\prod_i d_i\) and \(\prod_i e_i\) are squarefree and coprime to \(W\). Squarefreeness of \(\prod_i d_i\) forces each \(d_i\) to be squarefree and \((d_i, d_j) = 1\) for \(i \ne j\): a prime dividing two distinct coordinates would divide the product twice. The same holds for the \(e_i\).
Fix \(i\). Since \(d_i\) and \(e_i\) are squarefree, so is \([d_i, e_i]\) — indeed \([d_i, e_i] = (d_i, e_i) \cdot \frac{d_i}{(d_i, e_i)} \cdot \frac{e_i}{(d_i, e_i)}\) is a product of three pairwise coprime squarefree integers. For \(i \ne j\), a prime dividing both \([d_i, e_i]\) and \([d_j, e_j]\) would divide one of \(d_i, e_i\) and one of \(d_j, e_j\); each of the four cases contradicts one of \((d_i, d_j) = 1\), \((e_i, e_j) = 1\), \((d_i, e_j) = 1\), \((e_i, d_j) = 1\), the last two being part of the permissible-support condition. Hence \(\prod_i [d_i, e_i]\) is a product of pairwise coprime squarefree integers, so is squarefree.
Finally \(W\) is squarefree (it is a product of distinct primes) and coprime to \(\prod_i [d_i, e_i]\), since every prime factor of the latter divides some \(d_i\) or \(e_i\). A product of two coprime squarefree integers is squarefree, so \(q(\mathbf{d}, \mathbf{e}) = W \prod_i [d_i, e_i]\) is squarefree. ◻
Sublemma 1 (Cardinality of the \(q\)-fibre). Let \(W \ge 1\) and let \(r \ge 1\) be a multiple of \(W\). Then the set of pairs \((\mathbf{d}, \mathbf{e})\) with \(q(\mathbf{d}, \mathbf{e}) = r\) has cardinality at most \(\tau_{3k}(r/W)\).
Proof. We construct an injection from the fibre into the set of ordered factorisations of \(r/W\) into \(3k\) factors, whose cardinality is \(\tau_{3k}(r/W)\) by Definition 1.
Given a pair in the fibre, apply the per-coordinate lcm decomposition of Sublemma 1 in each coordinate: it sends the pair \((d_i, e_i)\) to the triple \(\bigl((d_i, e_i),\ d_i/(d_i, e_i),\ e_i/(d_i, e_i)\bigr)\), whose entries are pairwise coprime with product \([d_i, e_i]\), and it is a bijection onto such triples, with inverse \((g, a, b) \mapsto (g a, g b)\). Concatenating the \(k\) triples gives an ordered \(3k\)-tuple whose product is \(\prod_i [d_i, e_i] = q(\mathbf{d}, \mathbf{e})/W = r/W\).
The map is injective: from the \(3k\)-tuple one recovers each pair \((d_i, e_i)\) by the inverse of Sublemma 1, coordinate by coordinate. Injectivity into a set of cardinality \(\tau_{3k}(r/W)\) gives the bound. ◻
Proof uses: slem_pair_lcm_bijection
Sublemma 1 (Multiplicity of a squarefree modulus). Let \(\lambda\) have permissible support at \((R, W)\) with \(W \ge 1\), and let \(r \ge 1\) be squarefree. Then \[\bigl|\{(\mathbf{d}, \mathbf{e}) \text{ in the support of } \lambda : q(\mathbf{d}, \mathbf{e}) = r\}\bigr| \;\le\; \tau_{3k}(r).\]
Proof. If \(W \nmid r\) the fibre is empty, since \(W \mid q(\mathbf{d}, \mathbf{e})\) always. Otherwise Sublemma 1 bounds the fibre by \(\tau_{3k}(r/W)\). Since \(r\) is squarefree and \(W \mid r\), we have \(r = W \cdot (r/W)\) with \((W, r/W) = 1\); as \(\tau_{3k}\) is multiplicative and takes values \(\ge 1\) on positive integers, \(\tau_{3k}(r) = \tau_{3k}(W)\, \tau_{3k}(r/W) \ge \tau_{3k}(r/W)\). ◻
Proof uses: slem_S2m_q_factor_bijection
Sublemma 1 (Majorisation of the error sum by a sum over moduli). Let \(k \ge 2\), let \(\vartheta, \delta\) satisfy \(0 < \delta < \vartheta/2\), and let \(\lambda\) have permissible support at level \(\lfloor R \rfloor\) and modulus \(W\). Then there are a constant \(C > 0\) and a threshold \(N_0\), depending only on \(k\), \(\vartheta\) and \(\delta\), such that for all \(N \ge N_0\) \[\sum_{\mathbf{d}, \mathbf{e}} |\lambda(\mathbf{d})\, \lambda(\mathbf{e})|\, E\bigl(N, q(\mathbf{d}, \mathbf{e})\bigr) \;\le\; C\, \lambda_{\max}^2 \sum_{1 \le q \le N^{\vartheta}} \tau_{3k}(q)^2\, E(N, q),\] the outer sum being over pairs in the support of \(\lambda\).
Proof. Group the pairs according to the value \(r = q(\mathbf{d}, \mathbf{e})\). By Sublemma 1 every such \(r\) is squarefree, so the grouped sum is supported on squarefree \(r\), i.e. carries the factor \(\mu(r)^2\). By Sublemma 1 each value \(r\) is attained by at most \(\tau_{3k}(r)\) pairs, and \(|\lambda(\mathbf{d})\lambda(\mathbf{e})| \le \lambda_{\max}^2\) throughout. For the range: on the support of \(\lambda\) we have \(\prod_i d_i \le R\) and \(\prod_i e_i \le R\) (Definition 1), hence \(\prod_i [d_i, e_i] \le \prod_i d_i e_i \le R^2\) and \(r = W \prod_i [d_i, e_i] \le R^2 W\). ◻
Proof uses: slem_S2m_q_squarefree, slem_S2m_q_factor_count
Sublemma 1 (Divisor-weighted sum of local errors). Let \(k \ge 2\), let \(\vartheta \in (0, 1)\) be a level of distribution for the primes, and let \(A > 0\). Then there are \(C > 0\) and \(N_0\) such that for every \(N \ge N_0\), \[\sum_{1 \le q \le N^{\vartheta}} \tau_{3k}(q)^2\, E(N, q) \;\le\; C\, \frac{N}{(\log N)^{A}}.\]
Proof. Write \(E(N, q) = 1 + \mathcal{E}(N, q)\) with \(\mathcal{E} \ge 0\) the genuine discrepancy, and treat the two pieces separately.
The constant piece. A crude divisor bound suffices: for every \(\alpha > 0\) one has \(\tau_r(n) \ll_{r, \alpha} n^{\alpha}\), proved by comparing, at each prime power \(p^{a} \| n\), the polynomial factor \(\binom{a + r - 1}{r - 1}\) with the geometric factor \(p^{a\alpha}\), which dominates it for all \(p\) beyond a threshold depending on \(r\) and \(\alpha\), the finitely many remaining primes contributing a bounded factor. Taking \(\alpha\) so small that \(\vartheta(1 + 2\alpha) < 1\) gives \(\sum_{q \le N^{\vartheta}} \tau_{3k}(q)^2 \ll N^{\vartheta(1 + 2\alpha)} = o\bigl(N/(\log N)^{A}\bigr)\).
The discrepancy piece. By Cauchy–Schwarz, \[\sum_{q \le N^{\vartheta}} \tau_{3k}(q)^2\, \mathcal{E}(N, q) \;\le\; \Bigl(\sum_{q \le N^{\vartheta}} \frac{\tau_{3k}(q)^4}{\varphi(q)}\, N\Bigr)^{\!1/2} \Bigl(\sum_{q \le N^{\vartheta}} \frac{\varphi(q)\,\mathcal{E}(N, q)^2}{N}\Bigr)^{\!1/2}.\] For the first factor, split each \(q\) uniquely as a squarefree integer times a squarefull one: on squarefree \(q\) the sum \(\sum_{q \le z} \tau_{3k}(q)^4/\varphi(q) \ll_k z\,(\log z)^{(3k)^4}\) by a Mertens-type estimate (Lemma 1 applied with the parameter \((3k)^4\), using \(\tau_{3k}(q)^4 = \tau_{(3k)^4}(q)\) on squarefree arguments, which is Lemma 1 applied twice), while the squarefull part contributes a convergent factor because \(\sum_{w \text{ squarefull}} \tau_{3k}(w)^4/\varphi(w) < \infty\). Hence the first factor is \(\ll_k N^{(1 + \vartheta)/2} (\log N)^{B}\) for some \(B = B(k)\). For the second factor, the trivial bound \(\mathcal{E}(N, q) \ll N/\varphi(q)\) — a residue class modulo \(q\) meets \((N, 2N]\) in at most \(N/q + 1\) integers, and \(X_N \le N\) — gives \(\varphi(q)\mathcal{E}(N, q)/N \ll 1\), so the second factor is \(\ll \bigl(\sum_{q \le N^{\vartheta}} \mathcal{E}(N, q)\bigr)^{1/2}\). Finally the level of distribution \(\vartheta\) (Definition 1), in the dyadic form of Lemma 1, bounds \(\sum_{q \le N^{\vartheta}} \mathcal{E}(N, q)\) by \(\ll_{A'} N/(\log N)^{A'}\) for any prescribed \(A' > 0\).
Multiplying the two factors gives \(\ll_{k, A'} N/(\log N)^{A'/2 - B}\); choosing \(A' \ge 2A + 2B\) makes this \(\ll N/(\log N)^{A}\), and adding the constant piece completes the proof. ◻
Proof uses: lem_BV_restated, lem_mertens_tau_k, lem_tau_squared
Lemma 1 (Aggregate error bound). Let \(k \ge 2\), let \(A > 0\), and let \(\vartheta < 1/2\) be a level of distribution for the primes, with \(0 < \delta < \vartheta/2\). Then there are \(C > 0\) and \(N_0\) such that for every \(N \ge N_0\) and every \(\lambda\) with permissible support at \((R, W)\), \[\sum_{\mathbf{d}, \mathbf{e}} |\lambda(\mathbf{d})\,\lambda(\mathbf{e})|\, E\bigl(N, q(\mathbf{d}, \mathbf{e})\bigr) \;\le\; C\,\frac{y_{\max}^2\, N}{(\log N)^{A}}.\]
Uses: slem_S2m_triangle, slem_S2m_modulus_sum, lem_lambda_max_bound, def_adm_support, def_R, def_W_trick, not_error, def_q_lcm, thm_BV
Proof. By Sublemma 1 the left-hand side is at most \(\lambda_{\max}^2 \sum_{1 \le r \le R^2 W} \mu(r)^2 \tau_{3k}(r) E(N, r)\).
We first check that the range \(R^2 W\) lies inside the range \(N^{\vartheta}\) covered by the level of distribution. Indeed \(R^2 = N^{\vartheta - 2\delta}\) and \(W \le e^{2 D_0} = N^{o(1)}\) by Lemma 1, so \(R^2 W \le N^{\vartheta}\) for all \(N\) large enough, \(\delta > 0\) being fixed; the same computation, with the sharper bound of Lemma 1, places \(R^2 W\) below \(N^{1/2 - \eta}\) for some \(\eta > 0\), which is what makes Bombieri–Vinogradov applicable at all.
Bounding \(\mu(r)^2 \tau_{3k}(r) \le \tau_{3k}(r)^2\) and applying Sublemma 1 with the exponent \(A + 2k\) gives \[\sum_{1 \le r \le R^2 W} \mu(r)^2\, \tau_{3k}(r)\, E(N, r) \;\ll_{k, A}\; \frac{N}{(\log N)^{A + 2k}}.\] Finally Lemma 1 bounds the sieve coefficients in terms of the \(y\)-variables, \(\lambda_{\max} \ll_k y_{\max} (\log N)^{k}\), so \(\lambda_{\max}^2 \ll_k y_{\max}^2 (\log N)^{2k}\). Multiplying the two displays absorbs the \((\log N)^{2k}\) and leaves \(C\, y_{\max}^2 N/(\log N)^{A}\). ◻
Partial summation and the \(g\)-weighted Mertens sum
The \(s_{i,j}\)-estimates of the next subsection and the evaluation of the smooth case both rest on one arithmetic input: the Mertens-type sum \(\sum_{r < R,\, (r, W) = 1}\mu(r)^2 G(\log r/\log R)/g(r)\), and its multi-coordinate version. We record these here.
Two refinements of the one-variable partial-summation estimate of Lemma 1 are needed. First, iterating that estimate over several coordinates produces, at each intermediate stage, a profile obtained by integrating out the coordinates already processed; such a profile is bounded and Lipschitz, but there is no reason for it to be continuously differentiable, so a version of partial summation valid for Lipschitz profiles is required. Second, the accumulated error over \(k\) coordinates is acceptable only if the residual of the one-variable estimate is split into a piece decaying like a power of \(z\) and a piece of size \(O(1/\log z)\); the crude bound \(O(\log(2Vz)/\log z)\) of Lemma 1 is not enough to produce a \(\log D_0\) saving at each layer.
Sublemma 1 (Partial summation for a Lipschitz profile). There are functions \(C_1 = C_1(A_1, A_3) > 0\) and \(C_2 = C_2(A_1, A_3) > 0\) such that the following holds for every sieve datum \((\gamma, V, A_1, A_2, A_3, L)\), every real \(z \ge 2\), every \(M \ge 0\) and every \(G : [0,1] \to \mathbb{R}\) that is bounded by \(M\) and \(M\)-Lipschitz: \[\Bigl| \sum_{0 < d < z} \mu(d)^2 g_*(d)\, G\Bigl(\frac{\log d}{\log z}\Bigr) \;-\; \mathfrak{S}(\gamma)\,(\log z)\int_0^1 G \Bigr| \;\le\; 2M\, \mathcal{L}(\gamma, V, z),\] where \(\mathcal{L}(\gamma,V,z)\) denotes the real number satisfying \[\mathcal{L}(\gamma, V, z) \;=\; 2 C_1\, \mathfrak{S}(\gamma)\,(1 + \ell_V) \;+\; C_2\, \tau(V) \Bigl( z^{-1/8}\log(2Vz) \;+\; \frac{8\log(2V) + 64}{\log z} \Bigr).\]
Proof. For a \(C^1\) profile the estimate is the Abel-summation argument of Lemma 1, run without collapsing the two shapes of residual: writing \(H(t) := \sum_{d < t}\mu(d)^2 g_*(d)\) and \(H(t) = \mathfrak{S}(\gamma)\log t + r(t)\), Abel summation bounds the left-hand side by \(G_{\max}\bigl(|r(z)| + (\log z)^{-1}\int_1^z |r(t)|\,dt/t\bigr)\), and Sublemma 1 gives \(|r(t)| \le C_1\mathfrak{S}(\gamma)(1 + \ell_V) + C_2\tau(V)\,t^{-1/8}\log(2Vt)\). Integrating the first shape against \(dt/t\) contributes \(\mathfrak{S}(\gamma)(1 + \ell_V)\); integrating the second contributes \(\tau(V)\log(2V)/\log z\) up to absolute constants, since \(\int_1^\infty t^{-9/8}\log(2Vt)\,dt \ll \log(2V) + 1\); and the boundary term \(|r(z)|\) retains the shape \(\tau(V) z^{-1/8}\log(2Vz)\). This is the displayed bound with \(G_{\max}\) in place of \(2M\).
For a merely bounded Lipschitz \(G\), approximate: for each \(\eta > 0\) there is a \(C^1\) function \(G_\eta\) on \([0,1]\) with \(\sup|G - G_\eta| \le \eta\) and \(C^1\)-norm \(G_{\eta, \max} \le 2M\), obtained by convolving \(G\) with a smooth bump of width \(\eta/M\) (the Lipschitz bound controls the derivative of the mollification by \(M\), and the sup-norm by \(M\)). Both the discrete sum and the integral change by at most \(\eta\) times the total mass, which is bounded uniformly in \(\eta\); letting \(\eta \to 0\) transfers the \(C^1\) estimate to \(G\) with the constant \(2M\) in place of \(G_{\eta,\max}\). ◻
Proof uses: lem_partial_sum, slem_H_asymptotic
Lemma 1 (\(k\)-fold partial summation against a smooth profile). There are positive functions \(C=C(A_1,A_3)\), \(C_1=C_1(A_1,A_3)\) and \(C_2=C_2(A_1,A_3)\) such that the following holds for every sieve datum \((\gamma, V, A_1, A_2, A_3, L)\), every \(n \ge 1\), every real \(z \ge 2\) and every smooth \(F : \mathbb{R}^n \to \mathbb{R}\) supported in \(\mathcal{R}_n\): \[\Bigl| \sum_{\mathbf{u} \in \mathbb{N}^n} \Bigl(\prod_{i=1}^n \mu(u_i)^2 g_*(u_i)\Bigr) F\Bigl(\frac{\log u_1}{\log z}, \dotsc, \frac{\log u_n}{\log z}\Bigr)^{\!2} \;-\; \bigl(\mathfrak{S}(\gamma)\log z\bigr)^{n} \int_{\mathcal{R}_n} F^2 \Bigr| \;\le\; C\, n\, F_{\max}^2 \, B^{\,n-1}\, \mathcal{L}(\gamma, V, z),\] where \(\mathcal{L}(\gamma,V,z)\) and \(B\) are the real numbers satisfying \[\mathcal{L}(\gamma,V,z)=2C_1\mathfrak{S}(\gamma)(1+\ell_V) +C_2\tau(V)\left(z^{-1/8}\log(2Vz)+\frac{8\log(2V)+64}{\log z}\right)\] and \(B=2\mathfrak{S}(\gamma)\log z+\mathcal{L}(\gamma,V,z)\).
Proof. For each \(0 \le \ell \le n\), interpolate between the two sides through the \(n + 1\) layered evaluations \[\begin{aligned} E_\ell &\;:=\; \sum_{\mathbf{u} \in \mathbb{N}^{\ell}} \Bigl(\prod_{i \le \ell} \mu(u_i)^2 g_*(u_i)\Bigr) \bigl(\mathfrak{S}(\gamma)\log z\bigr)^{n - \ell}\\ &\qquad\qquad\times \int_{[0,1]^{n - \ell}} \mathbf{1}\Bigl[\textstyle\sum_{i \le \ell}\tfrac{\log u_i}{\log z} + \sum_{j > \ell} t_j \le 1\Bigr]\, F\Bigl(\tfrac{\log u_1}{\log z}, \dotsc, \tfrac{\log u_\ell}{\log z}, t_{\ell + 1}, \dotsc, t_n \Bigr)^{\!2} \, d\mathbf{t} \end{aligned}\] in which the first \(\ell\) coordinates are summed discretely against the sieve weight and the remaining \(n - \ell\) are integrated, each integrated coordinate paying the factor \(\mathfrak{S}(\gamma)\log z\) that partial summation would have produced. Since \(F\) is supported in \(\mathcal{R}_n\) the indicator may be dropped in \(E_0\), so \(E_0 = (\mathfrak{S}(\gamma)\log z)^{n}\int_{\mathcal{R}_n} F^2\); and \(E_n\) is the discrete sum on the left, the simplex indicator again being automatic from the support of \(F\).
One layer. Fix \(\ell\) and a prefix \(\mathbf{u} \in \mathbb{N}^{\ell}\), and let \(G_{\mathbf{u}}(s)\) denote the inner integral of \(F^2\) with the first \(\ell\) coordinates frozen at \(\log u_i/\log z\), the \((\ell+1)\)-st coordinate frozen at \(s\), and the remaining ones integrated over \([0,1]\) subject to the simplex constraint. Then \(|G_{\mathbf{u}}| \le F_{\max}^2\) pointwise, because the integrand is bounded by \(F_{\max}^2\) on a set of volume at most \(1\); and \(G_{\mathbf{u}}\) is \(2F_{\max}^2\)-Lipschitz in \(s\), because moving \(s\) moves both the integrand (with derivative bounded by \(2 F_{\max}^2\), as \(\partial_s F^2 = 2F\partial_s F\)) and the simplex constraint (whose symmetric difference has measure at most \(|s - s'|\), on which the integrand is bounded by \(F_{\max}^2\)). Applying Sublemma 1 with \(M = 2 F_{\max}^2\) replaces the discrete \((\ell+1)\)-st sum by \(\mathfrak{S}(\gamma)(\log z)\int_0^1 G_{\mathbf{u}}\), at a cost \(\le 4 F_{\max}^2 \mathcal{L}(\gamma, V, z)\) per prefix. Consequently \[|E_{\ell + 1} - E_{\ell}| \;\le\; \bigl(\mathfrak{S}(\gamma)\log z\bigr)^{n - \ell - 1} \Bigl(\sum_{\mathbf{u} \in \mathbb{N}^{\ell}} \prod_{i \le \ell}\mu(u_i)^2 g_*(u_i)\Bigr) \cdot 4 F_{\max}^2\, \mathcal{L}(\gamma, V, z),\] the prefix sum being restricted, by the support of \(F\), to \(u_i < z\).
The prefix mass. The prefix sum factorises over the \(\ell\) coordinates, and each single-coordinate sum \(\sum_{0 < d < z}\mu(d)^2 g_*(d)\) is, by Sublemma 1 with \(G \equiv 1\), at most \(\mathfrak{S}(\gamma)\log z + 2\mathcal{L}(\gamma, V, z) \le B\). Hence the prefix mass is at most \(B^{\ell}\), and since also \(\mathfrak{S}(\gamma)\log z \le B\) the whole bracket is at most \(B^{\,n-1}\).
Telescoping. Summing \(|E_{\ell+1} - E_{\ell}|\) over \(0 \le \ell < n\) gives \(|E_n - E_0| \le 4 n F_{\max}^2 B^{\,n-1}\mathcal{L}(\gamma, V, z)\), which is the claim. ◻
Proof uses: slem_partial_sum_lipschitz
Lemma 1 (The \(r_i\)-coordinate sum). Let \(\vartheta \in (0,1)\) and \(\delta \in (0, \vartheta/2)\). There are \(C > 0\) and \(N_0\) such that for every \(N \ge N_0\) with \(D_0 \ge 2\) and every \(G \in C^1([0,1])\), \[\Bigl| \sum_{\substack{r < R\\ (r, W) = 1}} \frac{\mu(r)^2}{g(r)}\, G\Bigl(\frac{\log r}{\log R}\Bigr) \;-\; \frac{\varphi(W)}{W}\,(\log R) \int_0^1 G \Bigr| \;\le\; C\,\frac{\varphi(W)}{W}\Bigl(\frac{\log R}{D_0}\Bigl|\int_0^1 G\Bigr| + (\log D_0)\, G_{\max}\Bigr).\]
Uses: slem_partial_sum_lipschitz, lem_maynard_sieve_datum, slem_gg_identification, slem_gg_singular_series, def_gamma_g, def_g_func, def_g_star, def_R, def_W_trick, lem_W_size
Proof. Apply the sharp partial-summation estimate of Sublemma 1 at \(z = R\) to the Maynard sieve datum built on the \(W\)-tricked density \(\gamma_g\) of Definition 1, namely the totally multiplicative function with \(\gamma_g(p) = 0\) for \(p \mid W\) and \(\gamma_g(p) = p/(p-1)\) for \(p \nmid W\). That this is a sieve datum, with \(V = W\), \(A_1 = 1/2\), \(A_3 = 2\) and \(L \ll \log D_0\), is Lemma 1; and by Sublemma 1 its derived multiplicative function is \(g_*(r) = \mu(r)^2/g(r)\) for \(r\) coprime to \(W\) and \(0\) otherwise, so the sum in the statement is exactly the sum that estimate governs.
By Sublemma 1, \(\mathfrak{S}(\gamma_g) = \frac{\varphi(W)}{W}\bigl(1 + O(1/D_0)\bigr)\). Substituting, the main term \(\mathfrak{S}(\gamma_g)(\log R)\int_0^1 G\) becomes \(\frac{\varphi(W)}{W}(\log R)\int_0^1 G\) together with an error \(O\bigl(\frac{\varphi(W)}{W}\frac{\log R}{D_0}\bigl|\int_0^1 G\bigr|\bigr)\) — the first of the two asserted error shapes, and the reason the statement carries \(\int_0^1 G\) rather than \(G_{\max}\) in that term.
For the residual \(\mathcal{L}(\gamma_g, W, R)\) of Sublemma 1: the term \(\mathfrak{S}(\gamma_g)(1 + \ell_W)\) is \(O\bigl(\frac{\varphi(W)}{W}\log D_0\bigr)\), since \(\ell_W = \sum_{p \le D_0}\frac{\log p}{p-1} = \log D_0 + O(1)\); and the term \(\tau(W)\bigl(R^{-1/8}\log(2WR) + (\log(2W)+1)/\log R\bigr)\) is negligible in comparison, because \(W = N^{o(1)}\) and \(\tau(W) = N^{o(1)}\) by Lemma 1 while \(\log R\) is a fixed positive multiple of \(\log N\). Multiplying by \(G_{\max}\) gives the second asserted error shape \(O\bigl(\frac{\varphi(W)}{W}(\log D_0)G_{\max}\bigr)\). ◻
The main term and the substitution \(\lambda \rightsquigarrow y^{(m)}\)
Definition 1 (The \(y^{(m)}\)-bilinear sum). For a finitely supported \(\lambda\) and an index \(m\), set \[\mathcal{Y}_m(\lambda) \;:=\; \sum_{\mathbf{u}} \Bigl(\prod_{i=1}^k \frac{\mu(u_i)^2}{g(u_i)}\Bigr) \ \sideset{}{^*}\sum_{s_{1,2}, \dotsc, s_{k, k-1}} \Bigl(\prod_{i \ne j} \frac{\mu(s_{i, j})}{g(s_{i, j})^2}\Bigr)\, y^{(m)}(\mathbf{a})\; y^{(m)}(\mathbf{b}),\] where \(\mathbf{a}\) and \(\mathbf{b}\) are the exponent tuples \(a_j = u_j \prod_{i \ne j} s_{j, i}\) and \(b_j = u_j \prod_{i \ne j} s_{i, j}\) of Definitions 1 and 1, and \(\sideset{}{^*}\sum\) is the restricted sum of Notation 1.
Uses: def_ym_vars, def_g_func, def_mobius, def_bold_a, def_bold_b, not_starsum
Lemma 1 (Main term of \(S_2^{(m)}(\lambda)\) after the Chinese remainder step). Let \(k \ge 2\), let \(\mathcal{H}\) have distinct shifts, and let \(\vartheta < 1/2\) be a level of distribution for the primes, with \(0 < \delta < \vartheta/2\). For every \(A > 0\) there are \(C, N_0 > 0\) such that for every \(N \ge N_0\), every compatible \(v_0\) and every \(\lambda\) with permissible support at \((R, W)\), \[\Bigl| S_2^{(m)}(\lambda) \;-\; \frac{X_N}{\varphi(W)} \sideset{}{'}\sum_{\substack{\mathbf{d}, \mathbf{e}\\ d_m = e_m = 1}} \frac{\lambda(\mathbf{d})\,\lambda(\mathbf{e})}{\prod_{i=1}^k \varphi([d_i, e_i])} \Bigr| \;\le\; C\,\frac{y_{\max}^2\, N}{(\log N)^{A}},\] where \(\sideset{}{'}\sum\) restricts to pairs with \(W, [d_1, e_1], \dotsc, [d_k, e_k]\) pairwise coprime and, in addition, \((h_m - h_i, [d_i, e_i]) = 1\) for every \(i \ne m\).
Uses: lem_S2m_expansion, lem_S2m_dm_em_one, lem_S2m_CRT, lem_S2m_error, def_S2m, def_adm_support, not_XN, def_totient, def_W_trick, thm_BV
Proof. By Lemma 1, \(S_2^{(m)}(\lambda) = \sum_{\mathbf{d}, \mathbf{e}} \lambda(\mathbf{d})\lambda(\mathbf{e})\, T_m(\mathbf{d}, \mathbf{e})\) with \(T_m(\mathbf{d}, \mathbf{e})\) the inner prime-weighted count.
Step 1: discard the pairs that cannot contribute. By Lemma 1, \(T_m(\mathbf{d}, \mathbf{e}) = 0\) for \(N\) large unless \(d_m = e_m = 1\). If two of \(W, [d_1, e_1], \dotsc, [d_k, e_k]\) share a prime \(p\), then the congruence system defining the inner sum is either inconsistent (so \(T_m = 0\)) or forces \(p \mid (h_i - h_j)\) for some \(i \ne j\); on the support of \(\lambda\) every prime factor of \([d_i, e_i]\) exceeds \(D_0\), and \(D_0 \to \infty\), so for \(N\) large this cannot happen for distinct shifts. The same argument shows the extra condition \((h_m - h_i, [d_i, e_i]) = 1\) is automatically satisfied for \(N\) large: any common prime would exceed \(D_0 > \max_{i \ne j}|h_i - h_j|\) and divide the nonzero integer \(h_m - h_i\). Thus, for \(N\) large, the outer sum may be restricted to \(\sideset{}{'}\sum\) with \(d_m = e_m = 1\), without changing its value.
Step 2: evaluate each retained inner sum. For a retained pair, all three hypotheses of Lemma 1 hold, so \(T_m(\mathbf{d}, \mathbf{e}) = X_N/\varphi(q(\mathbf{d}, \mathbf{e})) + O(E(N, q(\mathbf{d}, \mathbf{e})))\). Since \(W\) and the \([d_i, e_i]\) are pairwise coprime, multiplicativity of \(\varphi\) gives \(\varphi(q(\mathbf{d}, \mathbf{e})) = \varphi(W) \prod_i \varphi([d_i, e_i])\).
Step 3: split main and error. Substituting Step 2 into Step 1 produces exactly the displayed main term, plus an error at most a constant multiple of \(\sum_{\mathbf{d}, \mathbf{e}} |\lambda(\mathbf{d})\lambda(\mathbf{e})|\, E(N, q(\mathbf{d}, \mathbf{e}))\). Lemma 1 bounds this by \(C\, y_{\max}^2 N/(\log N)^{A}\). ◻
Proof uses: lem_S2m_expansion, lem_S2m_dm_em_one, lem_S2m_CRT, lem_S2m_error
Lemma 1 (\(\varphi\)-lcm decoupling). Let \(\lambda\) be supported on tuples whose coordinates are positive, squarefree and pairwise coprime. Then \[\sideset{}{'}\sum_{\substack{\mathbf{d}, \mathbf{e}\\ d_m = e_m = 1}} \frac{\lambda(\mathbf{d})\,\lambda(\mathbf{e})}{\prod_{i=1}^k \varphi([d_i, e_i])} \;=\; \sum_{\mathbf{u}} \Bigl(\prod_{i=1}^k g(u_i)\Bigr) \sideset{}{'}\sum_{\substack{\mathbf{d}, \mathbf{e}\\ u_i \mid (d_i, e_i)\ \forall i\\ d_m = e_m = 1}} \frac{\lambda(\mathbf{d})\,\lambda(\mathbf{e})}{\prod_{i=1}^k \varphi(d_i)\,\varphi(e_i)}.\]
Proof. For squarefree \(d, e\), Lemma 1 gives the divisor-sum identity \[\frac{1}{\varphi([d, e])} \;=\; \frac{1}{\varphi(d)\,\varphi(e)} \sum_{u \mid (d, e)} g(u),\] which is the multiplicative reformulation of \(\varphi([d, e])\,\varphi((d,e)) = \varphi(d)\varphi(e)\) together with \(\varphi= g * 1\) on squarefree arguments. Applying it in each of the \(k\) coordinates — legitimate, since the coordinates of tuples in the support of \(\lambda\) are squarefree — and expanding the resulting product of \(k\) finite divisor sums into a single sum over the tuple \(\mathbf{u} = (u_1, \dotsc, u_k)\) with \(u_i \mid (d_i, e_i)\) gives \[\prod_{i=1}^k \frac{1}{\varphi([d_i, e_i])} \;=\; \frac{1}{\prod_{i=1}^k \varphi(d_i)\varphi(e_i)} \sum_{\substack{\mathbf{u}\\ u_i \mid (d_i, e_i)\ \forall i}} \prod_{i=1}^k g(u_i).\] Multiplying by \(\lambda(\mathbf{d})\lambda(\mathbf{e})\), summing over the restricted set of pairs and interchanging the two finite summations yields the claim; the restriction \(\sideset{}{'}\sum\) and the pinning \(d_m = e_m = 1\) concern only the outer variables and are untouched. ◻
Proof uses: lem_lcm_phi_split, lem_phi_g_convolution
Lemma 1 (Möbius inversion of the cross-coprimality conditions). Let \(\lambda\) be supported on tuples whose coordinates are positive, squarefree and pairwise coprime. Then \[\sideset{}{'}\sum_{\substack{\mathbf{d}, \mathbf{e}\\ u_i \mid (d_i, e_i)\ \forall i\\ d_m = e_m = 1}} \frac{\lambda(\mathbf{d})\,\lambda(\mathbf{e})}{\prod_{i=1}^k \varphi(d_i)\,\varphi(e_i)} \;=\; \sum_{s_{1,2}, \dotsc, s_{k, k-1}} \Bigl(\prod_{i \ne j} \mu(s_{i, j})\Bigr) \sum_{\substack{\mathbf{d}, \mathbf{e}\\ u_i \mid (d_i, e_i)\ \forall i\\ s_{i, j} \mid (d_i, e_j)\ \forall i \ne j\\ d_m = e_m = 1}} \frac{\lambda(\mathbf{d})\,\lambda(\mathbf{e})}{\prod_{i=1}^k \varphi(d_i)\,\varphi(e_i)}.\]
Proof. By Lemma 1, which uses only the permissible-support hypothesis, the pairwise-coprimality restriction recorded by \(\sideset{}{'}\sum\) is, on the support of \(\lambda\), equivalent to the family of conditions \((d_i, e_j) = 1\) over ordered pairs \(i \ne j\): the conditions involving \(W\) and the same-index conditions hold automatically.
For each such ordered pair, Lemma 1 gives the Möbius detection identity \(\mathbf{1}[(d_i, e_j) = 1] = \sum_{s_{i, j} \mid (d_i, e_j)} \mu(s_{i, j})\), and \(s_{i, j} \mid (d_i, e_j)\) is the conjunction of \(s_{i, j} \mid d_i\) and \(s_{i, j} \mid e_j\). Taking the product of these \(k(k-1)\) identities and interchanging the resulting finite \((s_{i, j})\)-summation with the \((\mathbf{d}, \mathbf{e})\)-summation gives the displayed formula.
Moreover, exactly as in Lemma 1, every tuple \((s_{i,j})\) that contributes a nonzero term satisfies the coprimality restrictions recorded by \(\sideset{}{^*}\sum\): \((s_{i,j}, u_i) = (s_{i,j}, u_j) = 1\), \((s_{i,j}, s_{i,a}) = 1\) for \(a \ne j\) and \((s_{i,j}, s_{b,j}) = 1\) for \(b \ne i\). Each is a two-line contradiction with \((d_i, d_j) = 1\) or \((e_i, e_j) = 1\), both of which follow from squarefreeness of \(\prod_i d_i\) and of \(\prod_i e_i\). By Lemma 1 the restricted and unrestricted \(s\)-sums therefore agree. ◻
Lemma 1 (Substitution \(\lambda(\mathbf{d}) \rightsquigarrow y^{(m)}(\mathbf{r})\)). Let \(k \ge 2\), let \(\mathcal{H}\) have distinct shifts, and let \(\vartheta < 1/2\) be a level of distribution for the primes, with \(0 < \delta < \vartheta/2\). For every \(A > 0\) there are \(C, N_0 > 0\) such that for every \(N \ge N_0\), every compatible \(v_0\) and every \(\lambda\) with permissible support at \((R, W)\), \[\Bigl| S_2^{(m)}(\lambda) \;-\; \frac{X_N}{\varphi(W)}\, \mathcal{Y}_m(\lambda) \Bigr| \;\le\; C\, \frac{y_{\max}^2\, N}{(\log N)^{A}}.\]
Uses: lem_S2m_main_CRT, lem_S2m_decouple_phi_lcm, lem_S2m_mobius, def_ym_weighted_sum, def_ym_vars, def_g_func, not_XN, def_adm_support, thm_BV
Proof. By Lemma 1 it suffices to identify the restricted double sum, after the transformations of Lemmas 1 and 1, with \(\mathcal{Y}_m(\lambda)\); the error is carried unchanged. The identification is the \(S_2^{(m)}\)-analogue of the substitution performed for \(S_1\) (Lemma 1), the difference being that \(\lambda(\mathbf{d})\) is here normalised by \(\prod_i \varphi(d_i)\) rather than \(\prod_i d_i\), so that \(g\) replaces \(\varphi\) throughout.
Step 1: invert the definition of \(y^{(m)}\). Definition 1 reads \(y^{(m)}(\mathbf{r}) = \bigl(\prod_{i \ne m} \mu(r_i) g(r_i)\bigr) \sum_{\mathbf{d}:\ r_i \mid d_i,\ d_m = 1} \lambda(\mathbf{d})/\prod_{i \ne m} \varphi(d_i)\). Möbius inversion, coordinate by coordinate, gives on the support of \(\lambda\) \[\frac{\lambda(\mathbf{d})}{\prod_i \varphi(d_i)} \;=\; \Bigl(\prod_i \mu(d_i)\Bigr) \sum_{\substack{\mathbf{r}\\ d_i \mid r_i\ \forall i\\ r_m = 1}} \frac{y^{(m)}(\mathbf{r})}{\prod_i g(r_i)} \qquad (d_m = 1),\] the support constraint \(r_m = 1\) coming from Lemma 1.
Step 2: collapse the \(\mathbf{d}\)- and \(\mathbf{e}\)-sums. On the inner sum of Lemma 1 the constraints on \(\mathbf{d}\) are \(u_i \mid d_i\) and \(s_{i, j} \mid d_i\) for \(j \ne i\); by the coprimality of \(u_i\) with the \(s_{i, j}\) recorded in Lemma 1, these are equivalent to the single divisibility \(a_i \mid d_i\) with \(a_i = u_i \prod_{j \ne i} s_{i, j}\) (Definition 1). Inserting Step 1, swapping the finite \(\mathbf{d}\)- and \(\mathbf{r}\)-sums and applying Lemma 1 in the form \(\sum_{a_i \mid d_i \mid r_i} \mu(d_i) = \mu(a_i)\mathbf{1}[r_i = a_i]\) collapses the \(\mathbf{r}\)-sum to the single tuple \(\mathbf{r} = \mathbf{a}\), leaving \(\bigl(\prod_i \mu(a_i)\bigr) y^{(m)}(\mathbf{a})/\prod_i g(a_i)\). The identical computation on the \(\mathbf{e}\)-side — where the constraints are \(u_i \mid e_i\) and \(s_{j, i} \mid e_i\), with the two indices of \(s\) transposed — leaves \(\bigl(\prod_i \mu(b_i)\bigr) y^{(m)}(\mathbf{b})/\prod_i g(b_i)\) with \(b_i = u_i \prod_{j \ne i} s_{j, i}\) (Definition 1). The pinning \(d_m = e_m = 1\) forces \(u_m = 1\) and \(s_{m, j} = s_{j, m} = 1\), hence \(a_m = b_m = 1\).
Step 3: assemble the coefficient. The tuples \(\mathbf{a}, \mathbf{b}\) have squarefree entries, and \(\mu\) and \(g\) are multiplicative on them, so \(\mu(a_i) = \mu(u_i)\prod_{j \ne i}\mu(s_{i,j})\) and \(g(a_i) = g(u_i)\prod_{j \ne i} g(s_{i,j})\), and similarly for \(b_i\). Each \(\mu(s_{i,j})\) occurs once through \(\mu(a_i)\), once through \(\mu(b_j)\), and once through the inversion factor \(\prod_{i \ne j}\mu(s_{i,j})\) already present; since \(\mu^3 = \mu\), the net exponent is one. The \(g\)-factors combine as \(\bigl(\prod_i g(u_i)\bigr)\prod_i \bigl(g(a_i) g(b_i)\bigr)^{-1} = \prod_i g(u_i)^{-1} \prod_{i \ne j} g(s_{i,j})^{-2}\). The resulting coefficient is \(\bigl(\prod_i \mu(u_i)^2/g(u_i)\bigr) \prod_{i \ne j} \mu(s_{i,j})/g(s_{i,j})^2\), which is exactly the summand of Definition 1. ◻
Lemma 1 (Contribution of the terms with some \(s_{i, j} > D_0\)). Let \(\vartheta \in (0, 1)\) and \(\delta \in (0, \vartheta/2)\). There are \(C > 0\) and \(N_0\) such that for every \(N \ge N_0\), every \(\lambda\) with permissible support at \((R, W)\) and every pair of indices \(i \ne j\), the absolute contribution to \(\frac{X_N}{\varphi(W)}\mathcal{Y}_m(\lambda)\) of the terms with \(s_{i, j} > D_0\) is at most \[C\, \frac{(y^{(m)}_{\max})^2\, \varphi(W)^{k-2}\, N\, (\log R)^{k-1}}{W^{k-1}\, D_0\, \log N}.\]
Uses: def_ym_weighted_sum, def_y_max, def_g_func, def_D_0, def_W_trick, def_R, not_XN
Proof. The argument is the \(g\)-analogue of Lemma 1, with three substitutions: the prefactor \(N/W\) becomes \(X_N/\varphi(W)\), the multiplicative weight \(\varphi\) becomes \(g\), and the convergent series of Lemma 1 becomes that of Lemma 1.
Fix the offending pair \((i, j)\). Bound \(|y^{(m)}(\mathbf{a})\, y^{(m)}(\mathbf{b})|\) by \((y^{(m)}_{\max})^2\), which is finite by Lemma 1, and factor the remaining sum into independent coordinate sums. There are three kinds of factor.
(a) The \(k - 1\) free \(u\)-coordinates. Since \(u_m = 1\) on the support, the \(u\)-sum contributes \(\bigl(\sum_{u \le R,\ (u, W) = 1} \mu(u)^2/g(u)\bigr)^{k-1}\). The \(g\)-analogue of the Mertens estimate of Lemma 1 — available here as the \(G \equiv 1\) case of Lemma 1, or directly by comparing \(1/g(p)\) with \(1/\varphi(p)\), whose difference \(1/((p-1)(p-2)) \ll p^{-2}\) sums to \(O(1/D_0)\) over \(p > D_0\) — bounds this by \(\ll_k \varphi(W)^{k-1}(\log R)^{k-1}/W^{k-1}\).
(b) The distinguished big coordinate. By Lemmas 1 and 1, every contributing \(s_{i, j} \ne 1\) is coprime to \(W\) and has all its prime factors \(> D_0\). For squarefree \(s\) with least prime factor \(> D_0 \ge 9\) one has \(g(p) = p - 2 \ge p/3\), hence \(1/g(s)^2 \le \prod_{p \mid s} 9/p^2\), and summing the resulting multinomial series gives \[\sum_{\substack{s > 1,\ \mu(s)^2 = 1\\ p \mid s \Rightarrow p > D_0}} \frac{1}{g(s)^2} \;\le\; \exp(9/D_0) - 1 \;\ll\; \frac{1}{D_0}.\]
(c) The remaining \(k^2 - k - 1\) coordinates \(s_{a, b}\). Each contributes \(\sum_{s \ge 1} \mu(s)^2/g(s)^2 \ll 1\), convergent by Lemma 1.
Multiplying (a), (b), (c) by the prefactor and using \(X_N \ll N/\log N\) (a consequence of Theorem 3, or of Chebyshev’s bound) gives the stated estimate. ◻
Lemma 1 (Contribution of the diagonal \(s_{i, j} = 1\)). For every \(N\), \(W\) and every finitely supported \(\lambda\), the contribution to \(\frac{X_N}{\varphi(W)}\, \mathcal{Y}_m(\lambda)\) from the single tuple with \(s_{i, j} = 1\) for all \(i \ne j\) is \[\frac{X_N}{\varphi(W)} \sum_{\mathbf{u}} \frac{\bigl(y^{(m)}(\mathbf{u})\bigr)^2}{\prod_{i=1}^k g(u_i)}.\]
Proof. Setting \(s_{i, j} = 1\) for all \(i \ne j\) in \(a_j = u_j \prod_{i \ne j} s_{j, i}\) and \(b_j = u_j \prod_{i \ne j} s_{i, j}\) gives \(\mathbf{a} = \mathbf{b} = \mathbf{u}\), so the product \(y^{(m)}(\mathbf{a})\,y^{(m)}(\mathbf{b})\) becomes \((y^{(m)}(\mathbf{u}))^2\). The \(s\)-coefficient is \(\prod_{i \ne j} \mu(1)/g(1)^2 = 1\) since \(g(1) = 1\). Finally \(y^{(m)}(\mathbf{u})\) vanishes unless every \(u_i\) is squarefree (Lemma 1), so on the support \(\mu(u_i)^2 = 1\) and the \(u\)-coefficient \(\prod_i \mu(u_i)^2/g(u_i)\) reduces to \(1/\prod_i g(u_i)\). ◻
Proof uses: lem_S1_sij_one
Lemma 1 (Absorption of the \(X_N\) error). Let \(k \ge 2\), \(\vartheta \in (0, 1)\) and \(\delta \in (0, \vartheta/2)\). There are \(C > 0\) and \(N_0\) such that for every \(N \ge N_0\) and every \(\lambda\) with permissible support at \((R, W)\), \[\Bigl| \Bigl(\frac{X_N}{\varphi(W)} - \frac{N}{\varphi(W)\log N}\Bigr) \sum_{\mathbf{u}} \frac{(y^{(m)}(\mathbf{u}))^2}{\prod_{i=1}^k g(u_i)} \Bigr| \;\le\; C\,\frac{(y^{(m)}_{\max})^2\, \varphi(W)^{k-2}\, N\, (\log N)^{k-2}}{W^{k-1}\, D_0}.\]
Uses: thm_PNT_XN, not_XN, def_ym_vars, def_y_max, def_g_func, def_R, def_D_0, def_W_trick
Proof. By Theorem 3, \(X_N = N/\log N + O(N/(\log N)^2)\), so the prefactor difference is \(O\bigl(N/(\varphi(W)(\log N)^2)\bigr)\).
It remains to bound the diagonal sum. Bounding \((y^{(m)}(\mathbf{u}))^2\) by \((y^{(m)}_{\max})^2\) and using \(u_m = 1\) together with the coordinate restrictions \(u_i \le R\), \((u_i, W) = 1\) and \(\mu(u_i)^2 = 1\) from Lemma 1, the sum factorises across the \(k - 1\) free coordinates: \[\sum_{\mathbf{u}} \frac{(y^{(m)}(\mathbf{u}))^2}{\prod_i g(u_i)} \;\le\; (y^{(m)}_{\max})^2 \Bigl(\sum_{\substack{u \le R\\ (u, W) = 1}} \frac{\mu(u)^2}{g(u)}\Bigr)^{\!k-1} \;\ll_k\; \frac{(y^{(m)}_{\max})^2\, \varphi(W)^{k-1}\, (\log R)^{k-1}}{W^{k-1}},\] the last step being the \(G \equiv 1\) case of Lemma 1. Multiplying the two displays and using \(\log R \le \log N\) gives a bound \(\ll (y^{(m)}_{\max})^2 \varphi(W)^{k-2} N (\log N)^{k-3}/W^{k-1}\), which is smaller than the asserted bound because \((\log N)^{k-3} \, D_0 \le (\log N)^{k-2}\) for \(N\) large: \(D_0\) grows only triple-logarithmically. ◻
Proof uses: thm_PNT_XN, lem_S2m_g_coord, lem_ym_support, lem_ym_max_finite
Lemma 1 (The second moment in terms of \(y^{(m)}\)). Let \(k \ge 2\), let \(\mathcal{H}\) have distinct shifts, let \(\vartheta < 1/2\) be a level of distribution for the primes and \(0 < \delta < \vartheta/2\). For every \(A > 0\) there are \(C > 0\) and \(N_0\) such that for every \(N \ge N_0\), every compatible \(v_0\) and every \(\lambda\) with permissible support at \((R, W)\), \[\begin{gathered} \Bigl| S_2^{(m)}(\lambda) \;-\; \frac{N}{\varphi(W)\,\log N} \sum_{\mathbf{r}} \frac{\bigl(y^{(m)}(\mathbf{r})\bigr)^2}{\prod_{i=1}^k g(r_i)} \Bigr|\\ \;\le\; C\,\frac{(y^{(m)}_{\max})^2\, \varphi(W)^{k-2}\, N\, (\log N)^{k-2}}{W^{k-1}\, D_0} \;+\; C\,\frac{y_{\max}^2\, N}{(\log N)^{A}}. \end{gathered}\]
Uses: lem_S2m_substitute_ym, lem_S2m_sij_D0, lem_S2m_sij_one, lem_S2m_PNT, def_S2m, def_ym_vars, def_y_max, def_g_func, def_D_0, def_W_trick, thm_BV, thm_PNT_XN
Proof. By Lemma 1, \(S_2^{(m)}(\lambda)\) agrees with \(\frac{X_N}{\varphi(W)}\mathcal{Y}_m(\lambda)\) up to \(C\, y_{\max}^2 N/(\log N)^{A}\), which is the second error term of the claim.
Split the \(s\)-summation in \(\mathcal{Y}_m(\lambda)\) into the diagonal tuple, all \(s_{i, j} = 1\), and its complement. By Lemma 1 — whose proof uses only that \(\lambda\) is supported on tuples \(\mathbf{d}\) with \(\prod_i d_i\) squarefree and coprime to \(W\), and that \(s_{i, j} \mid d_i\), both of which hold here — every contributing \(s_{i, j}\) is either \(1\) or has all prime factors \(> D_0\); in particular any off-diagonal tuple has some \(s_{i, j} > D_0\). Summing the bound of Lemma 1 over the \(k(k-1)\) ordered pairs (Lemma 1) shows the off-diagonal contribution is \[\ll_k \frac{(y^{(m)}_{\max})^2\, \varphi(W)^{k-2}\, N\, (\log R)^{k-1}}{W^{k-1}\, D_0\, \log N} \;\le\; \frac{(y^{(m)}_{\max})^2\, \varphi(W)^{k-2}\, N\, (\log N)^{k-2}}{W^{k-1}\, D_0},\] using \(\log R \le \log N\).
The diagonal contribution is evaluated by Lemma 1 as \(\frac{X_N}{\varphi(W)}\sum_{\mathbf{r}} (y^{(m)}(\mathbf{r}))^2/\prod_i g(r_i)\), and Lemma 1 replaces \(X_N\) by \(N/\log N\) at a cost within the same first error term. Adding the three contributions gives the claim. ◻
The smooth case
Lemma 1 reduces \(S_2^{(m)}(\lambda)\) to the diagonal quadratic form \(\sum_{\mathbf{r}} (y^{(m)}(\mathbf{r}))^2/\prod_{i \ne m} g(r_i)\) — the \(i = m\) factor being \(g(1) = 1\) because \(y^{(m)}\) is supported on \(r_m = 1\). We now evaluate that form for the \(F\)-derived weight \(\lambda_0\) of Definition 1.
The evaluation does not go through a per-tuple asymptotic for \(y^{(m)}(\mathbf{r})\). Up to a negligible error, \(y^{(m)}(\mathbf{r})\) equals the inverse sum \(Y^{(m)}(\mathbf{r})\) of Definition 1, whose defining coprimality condition carries the large modulus \(W \prod_{i \ne m} r_i\); a per-tuple evaluation would need partial summation at that modulus, which is not available. Instead we square first and sum second: the large modulus then appears as a coupling between the \(k + 1\) summation coordinates \(u, u', (r_i)_{i \ne m}\), and every such coupling can be dropped at a relative cost \(O(1/D_0)\), exactly as in the \(S_1\) analysis, because two coordinates coprime to \(W\) can only share a prime \(> D_0\). Once the couplings are dropped, each coordinate is coprime to \(W\) alone, and the \((k+1)\)-fold sum is evaluated coordinate by coordinate at modulus \(W\).
Definition 1 (The decoupled \((k+1)\)-fold sum). Let \(F : [0, 1]^k \to \mathbb{R}\). For an index \(m\) set \[\mathcal{D}_m(F) \;:=\; \sum_{\substack{u, u',\, (\rho_i)_{i \ne m}\\ \text{all squarefree and coprime to } W}} \frac{\mu(u)^2\, \mu(u')^2}{\varphi(u)\,\varphi(u')} \Bigl(\prod_{i \ne m} \frac{\mu(\rho_i)^2}{g(\rho_i)}\Bigr)\, \Phi\Bigl(\bigl(\tfrac{\log \rho_i}{\log R}\bigr)_{i \ne m};\ \tfrac{\log u}{\log R},\ \tfrac{\log u'}{\log R}\Bigr),\] where \(\Phi\bigl((t_i)_{i \ne m}; s, s'\bigr) := F(\dotsc, s, \dotsc)\, F(\dotsc, s', \dotsc)\) places \(s\), respectively \(s'\), in the \(m\)-th slot of two copies of \(F\) and \(t_i\) in the others. No coprimality is imposed between the summation variables.
Uses: def_W_trick, def_R, def_mobius, def_totient, def_g_func
Lemma 1 (Reduction of the diagonal form to a second moment of \(Y\)). Let \(k \ge 2\) and let \(F : [0,1]^k \to \mathbb{R}\) be smooth and supported in \(\mathcal{R}_k\). There is a constant \(C \ge 0\), depending only on \(k\), and a threshold \(N_0\) depending on \(F, \vartheta, \delta\), such that for every \(N \ge N_0\), every \(\lambda\) with permissible support at \((R, W)\) and \(y_{\max} \le F_{\max}\), and every \(m\), \[\Bigl| \sum_{\mathbf{r}} \frac{(y^{(m)}(\mathbf{r}))^2}{\prod_{i \ne m} g(r_i)} \;-\; \sum_{\mathbf{r}} \frac{Y^{(m)}(\mathbf{r})^2}{\prod_{i \ne m} g(r_i)} \Bigr| \;\le\; C\, \frac{F_{\max}^2\, \varphi(W)^{k+1}\, (\log R)^{k+1}}{W^{k+1}\, D_0},\] both sums being over tuples \(\mathbf{r}\) with \(r_m = 1\) and \(r_i \ge 1\) coprime to \(W\) for \(i \ne m\).
Uses: def_ym_marginal, lem_ym_substitute_smooth, lem_S2m_g_coord, def_ym_vars, def_g_func, def_F_smooth_max, def_y_max, def_adm_support, def_R, def_W_trick, def_simplex
Proof. By Lemma 1, \(y^{(m)}(\mathbf{r}) = Y^{(m)}(\mathbf{r}) + \delta_{\mathbf{r}}\) with a uniform bound \(|\delta_{\mathbf{r}}| \le \eta := c\, F_{\max}\,\varphi(W)(\log R)/(W D_0)\): the passage from \(y^{(m)}\) to the inverse sum costs only the terms in which the auxiliary tuple differs from \(\mathbf{r}\) in some coordinate \(j \ne m\), and those are controlled by a \(1/D_0\) saving of the same shape as in the \(S_1\) analysis.
Next, \(|Y^{(m)}(\mathbf{r})| \le F_{\max} \sum_{(a, W) = 1} \mu(a)^2/\varphi(a) \ll F_{\max}\,\varphi(W)(\log R)/W\) by the Mertens estimate of Lemma 1, and hence also \(|y^{(m)}(\mathbf{r})| \ll F_{\max}\varphi(W)(\log R)/W\). Therefore \[(y^{(m)}(\mathbf{r}))^2 - Y^{(m)}(\mathbf{r})^2 = 2 Y^{(m)}(\mathbf{r})\delta_{\mathbf{r}} + \delta_{\mathbf{r}}^2 = O\bigl(\eta\, F_{\max}\varphi(W)(\log R)/W\bigr).\] Summing this against \(1/\prod_{i \ne m} g(r_i)\) and using \[\sum_{\mathbf{r}} \prod_{i \ne m} \frac{\mu(r_i)^2}{g(r_i)} \;\le\; \prod_{i \ne m} \sum_{\substack{r < R\\ (r, W) = 1}} \frac{\mu(r)^2}{g(r)} \;\ll_k\; \Bigl(\frac{\varphi(W)\log R}{W}\Bigr)^{\!k-1}\] — the case \(G \equiv 1\) of Lemma 1 — gives the total \(\ll \eta\, F_{\max}\,\varphi(W)^{k}(\log R)^{k}/W^{k} = F_{\max}^2 \varphi(W)^{k+1}(\log R)^{k+1}/(W^{k+1} D_0)\), as claimed. ◻
Proof uses: lem_ym_substitute_smooth, lem_mertens_W, lem_S2m_g_coord
Lemma 1 (Expanding the square and dropping the couplings). Let \(k \ge 2\) and let \(F : [0,1]^k \to \mathbb{R}\) be smooth and supported in \(\mathcal{R}_k\), and let \(\lambda_0\) be the associated \(F\)-derived weight. There is a constant \(C \ge 0\) depending only on \(k\), and a threshold \(N_0\) depending on \(F, \vartheta, \delta\), such that for every \(N \ge N_0\) and every \(m\), \[\Bigl| \sum_{\mathbf{r}} \frac{Y^{(m)}(\mathbf{r})^2}{\prod_{i \ne m} g(r_i)} \;-\; \mathcal{D}_m(F) \Bigr| \;\le\; C\, \frac{F_{\max}^2\, \varphi(W)^{k+1}\, (\log R)^{k+1}}{W^{k+1}\, D_0}.\]
Uses: def_ym_marginal, def_decoupled_sum, def_lambda_from_F, lem_smooth_y, lem_S2m_g_coord, def_g_func, def_F_smooth_max, def_R, def_W_trick, def_simplex
Proof. For the \(F\)-derived weight, Lemma 1 evaluates the \(y\)-variables in closed form, and \(Y^{(m)}(\mathbf{r})\) becomes \[Y^{(m)}(\mathbf{r}) \;=\; \sum_{\substack{u \ge 1\\ (u,\ W \prod_{i \ne m} r_i) = 1}} \frac{\mu(u)^2}{\varphi(u)}\, F\Bigl(\dotsc, \tfrac{\log u}{\log R}, \dotsc\Bigr).\]
Step 1: square. The condition \((u, W\prod_{i \ne m} r_i) = 1\) is the conjunction of \((u, W) = 1\) and \((u, r_i) = 1\) for \(i \ne m\). Since each \(r_i\) is squarefree, \(\mathbf{1}[(u, r_i) = 1]\,\mathbf{1}[(u', r_i) = 1] = \mathbf{1}[(u u', r_i) = 1]\), so \[Y^{(m)}(\mathbf{r})^2 = \sum_{\substack{u, u'\\ (u u', W) = 1}} \frac{\mu(u)^2\mu(u')^2}{\varphi(u)\varphi(u')} \Bigl[\prod_{i \ne m}\mathbf{1}[(u u', r_i) = 1]\Bigr]\,\Phi.\] All variables are \(\le R\) on the support of \(F\), so summing over \(\mathbf{r}\) and interchanging the finite sums is legitimate. The result is \(\mathcal{D}_m(F)\) with the couplings \((u u', r_i) = 1\) (\(i \ne m\)) and \((r_i, r_j) = 1\) (\(i \ne j\)) still imposed.
Step 2: drop the couplings. All of \(u, u', r_i\) are coprime to \(W\), so a prime \(p\) shared by two of them satisfies \(p > D_0\). On squarefree arguments the local weight at such a prime is \(1/\varphi(p) = 1/(p-1)\) for the \(u\)-coordinates and \(1/g(p) = 1/(p-2)\) for the \(r\)-coordinates, so restoring a forbidden pair costs a factor at most \(1/((p-1)(p-2))\) or \(1/(p-2)^2\), and \(\sum_{p > D_0} (p-2)^{-2} \ll 1/D_0\). Exactly as in Lemma 1, summing over the \(O(k^2)\) pairs bounds the total cost of removing all couplings by \(O(1/D_0)\) times the fully unrestricted sum with \(|\Phi|\) in place of \(\Phi\).
Step 3: size of the unrestricted sum. Using \(|\Phi| \le F_{\max}^2\), the Mertens estimate of Lemma 1 for the two \(u\)-coordinates and Lemma 1 with \(G \equiv 1\) for the \(k-1\) coordinates \(r_i\), \[\sum_{\substack{u, u', (r_i)_{i \ne m}\\ \text{coprime to } W}} \frac{\mu(u)^2\mu(u')^2}{\varphi(u)\varphi(u')}\Bigl(\prod_{i \ne m}\frac{\mu(r_i)^2}{g(r_i)}\Bigr) |\Phi| \;\ll_k\; F_{\max}^2 \Bigl(\frac{\varphi(W)\log R}{W}\Bigr)^{\!k+1}.\] Combining with Step 2 gives the stated error. ◻
Proof uses: lem_smooth_y, lem_S1_drop_uij, lem_mertens_W, lem_S2m_g_coord
Lemma 1 (Evaluation of the decoupled \((k+1)\)-fold sum). Let \(k \ge 2\). There is a constant \(C \ge 0\), depending only on \(k\), such that for every smooth \(F : [0,1]^k \to \mathbb{R}\) supported in \(\mathcal{R}_k\) there is a threshold \(N_0\) with \[\Bigl| \mathcal{D}_m(F) \;-\; \frac{\varphi(W)^{k+1}\,(\log R)^{k+1}}{W^{k+1}}\, J_k^{(m)}(F) \Bigr| \;\le\; C\,\frac{F_{\max}^2\,\varphi(W)^{k+1}\,(\log R)^{k+1}}{W^{k+1}\, D_0}\] for every \(N \ge N_0\) and every \(m\).
Uses: def_decoupled_sum, def_J_k, lem_kfold_partial_sum, slem_partial_sum_lipschitz, lem_S2m_g_coord, slem_gg_identification, slem_gg_singular_series, def_F_smooth_max, def_R, def_W_trick, def_simplex
Proof. Separate the two \(m\)-slot coordinates \(u, u'\), whose weight is \(\mu^2/\varphi\), from the \(k - 1\) outer coordinates \((\rho_i)_{i \ne m}\), whose weight is \(\mu^2/g\).
Step 1: the two \(m\)-slot coordinates. With the outer coordinates frozen, the \(u\)-sum is \(\sum_{u < R,\ (u, W) = 1} \frac{\mu(u)^2}{\varphi(u)} F\bigl(\dotsc, \tfrac{\log u}{\log R}, \dotsc\bigr)\), a partial sum for the datum \(\gamma_W\) with \(\mathfrak{S}(\gamma_W) = \varphi(W)/W\); the profile is \(F\) with the \(m\)-th slot free, whose \(C^1\)-norm is at most \(F_{\max}\). One application of Sublemma 1 replaces it by \(\frac{\varphi(W)}{W}(\log R)\, F^{(m)}\bigl((t_i)_{i \ne m}\bigr)\), where \(F^{(m)}\bigl((t_i)_{i \ne m}\bigr) = \int_0^1 F(\dotsc, t_m, \dotsc)\, dt_m\) is the marginal of \(F\) in the \(m\)-th coordinate, with error \(O\bigl(\frac{\varphi(W)}{W}(\log D_0)F_{\max}\bigr)\); the residual of Sublemma 1 is of that size because \(V = W = N^{o(1)}\) while \(\log R\) is a fixed positive multiple of \(\log N\). Doing the same for \(u'\) and multiplying, the two \(m\)-slot coordinates contribute the factor \(\bigl(\frac{\varphi(W)}{W}\log R\bigr)^{2}\) and replace the coupling \(\Phi\) by \(\bigl(F^{(m)}\bigr)^2\), with a relative error \(O(\log D_0/\log R) = O(1/D_0)\).
Step 2: the \(k-1\) outer coordinates. What remains is the \((k-1)\)-fold sum \[\sum_{(\rho_i)_{i \ne m}} \Bigl(\prod_{i \ne m} \mu(\rho_i)^2\, g_*(\rho_i)\Bigr) \, F^{(m)}\Bigl(\bigl(\tfrac{\log \rho_i}{\log R}\bigr)_{i \ne m}\Bigr)^{\!2},\] where \(g_*\) is the derived function of the datum \(\gamma_g\), equal to \(\mu^2/g\) on integers coprime to \(W\) and \(0\) otherwise by Sublemma 1. The marginal \(F^{(m)}\) is smooth and supported in \(\mathcal{R}_k^{(m)}\), the \((k-1)\)-dimensional simplex, so Lemma 1 applies with \(n = k - 1\), \(z = R\) and profile \(F^{(m)}\), giving \[\bigl(\mathfrak{S}(\gamma_g)\log R\bigr)^{k-1}\int_{\mathcal{R}_k^{(m)}} \bigl(F^{(m)}\bigr)^2 \;+\; O\Bigl(F_{\max}^2 \Bigl(\frac{\varphi(W)\log R}{W}\Bigr)^{\!k-2} \frac{\varphi(W)}{W}\log D_0 \Bigr),\] again using \(V = W = N^{o(1)}\) to reduce the residual \(\mathcal{L}(\gamma_g, W, R)\) to \(O\bigl(\frac{\varphi(W)}{W}\log D_0\bigr)\) and \(B\) to \(O\bigl(\frac{\varphi(W)}{W}\log R\bigr)\). By Definition 1, \(\int_{\mathcal{R}_k^{(m)}} (F^{(m)})^2 = J_k^{(m)}(F)\).
Step 3: assembly. By Sublemma 1, \(\mathfrak{S}(\gamma_g) = \frac{\varphi(W)}{W}\bigl(1 + O(1/D_0)\bigr)\), and for fixed \(k\), \((1 + O(1/D_0))^{k-1} = 1 + O(1/D_0)\). Multiplying Steps 1 and 2 therefore gives the main term \(\frac{\varphi(W)^{k+1}(\log R)^{k+1}}{W^{k+1}}J_k^{(m)}(F)\). Since \(J_k^{(m)}(F) \le F_{\max}^2\), the deviation of the singular-product factor from \(1\) costs \(O\bigl(F_{\max}^2\varphi(W)^{k+1}(\log R)^{k+1}/(W^{k+1}D_0)\bigr)\). Each remaining error is a partial-summation residual at one coordinate multiplied by the leading factors at the other \(k\), hence of size \(O\bigl(F_{\max}^2\varphi(W)^{k+1}(\log D_0)(\log R)^{k}/W^{k+1}\bigr)\), whose ratio to the main term is \(O(\log D_0/\log R) = O(1/D_0)\). Collecting the errors gives the claim. ◻
Lemma 1 (Asymptotic for \(S_2^{(m)}\) at the smooth weight). Let \(k \ge 2\), let \(\mathcal{H}\) have distinct shifts, let \(F : [0,1]^k \to \mathbb{R}\) be smooth and supported in \(\mathcal{R}_k\), and let \(\lambda_0\) be the associated \(F\)-derived weight. Let \(\vartheta < 1/2\) be a level of distribution for the primes and \(0 < \delta < \vartheta/2\). Then there are \(C > 0\) and \(N_0\) such that for every \(N \ge N_0\) and every compatible \(v_0\), \[\Bigl| S_2^{(m)}(\lambda_0) \;-\; \frac{\varphi(W)^{k}\, N\, (\log R)^{k+1}}{W^{k+1}\,\log N}\, J_k^{(m)}(F) \Bigr| \;\le\; C\, \frac{F_{\max}^2\, \varphi(W)^{k}\, N\, (\log R)^{k}}{W^{k+1}\, D_0}.\]
Uses: lem_S2m_from_ym, lem_S2m_second_moment, lem_S2m_expand_drop, lem_S2m_eval, def_lambda_from_F, def_J_k, def_S2m, def_F_smooth_max, def_R, def_W_trick, def_D_0, def_simplex, thm_BV, thm_PNT_XN
Proof. Apply Lemma 1 to \(\lambda = \lambda_0\) with the exponent \(A = 2k + 3\), and then chain Lemmas 1, 1 and 1 to evaluate the diagonal form \(\sum_{\mathbf{r}} (y^{(m)}(\mathbf{r}))^2/\prod_{i \ne m} g(r_i)\). The three chained lemmas each cost \(O\bigl(F_{\max}^2 \varphi(W)^{k+1}(\log R)^{k+1}/(W^{k+1}D_0)\bigr)\), so their sum does too, and the main term is \(\frac{\varphi(W)^{k+1}(\log R)^{k+1}}{W^{k+1}}J_k^{(m)}(F)\). Multiplying by the prefactor \(N/(\varphi(W)\log N)\) from Lemma 1 gives the asserted main term and an error \[O\Bigl(\frac{F_{\max}^2\,\varphi(W)^{k}\, N\,(\log R)^{k+1}}{W^{k+1}\, D_0\, \log N}\Bigr) \;=\; O\Bigl(\frac{F_{\max}^2\,\varphi(W)^{k}\, N\,(\log R)^{k}}{W^{k+1}\, D_0}\Bigr),\] using \(\log R \le \log N\).
Two further errors must be absorbed into the same shape. The first error term of Lemma 1 is \((y^{(m)}_{\max})^2 \varphi(W)^{k-2} N (\log N)^{k-2}/(W^{k-1}D_0)\); for the \(F\)-derived weight \(y^{(m)}_{\max} \ll_k F_{\max}\,\varphi(W)(\log R)/W\) by Lemma 1 applied to the defining sum of \(y^{(m)}\), and \(\log N \le (\vartheta/2 - \delta)^{-1}\log R\), so this term is again \(O\bigl(F_{\max}^2\varphi(W)^{k}N(\log R)^{k}/(W^{k+1}D_0)\bigr)\). The second error term is \(y_{\max}^2 N/(\log N)^{A}\) with \(y_{\max} \le F_{\max}\); since \(\varphi(W)^{k}(\log R)^{k}/(W^{k+1}D_0) \gg (\log N)^{-(2k+3)}\) for large \(N\) — the left side decays only like a power of \(\log N\) times \(1/(W D_0)\), and \(W = N^{o(1)}\) — the choice \(A = 2k + 3\) absorbs it as well. ◻
The two-weight form and the outer support split
The preceding subsections treat \(S_2^{(m)}(\lambda)\) as a quadratic form in a single weight. Two further pieces of the reduction are bilinear, and are recorded here in the two-weight form in which they are proved and used.
The first is that \(\lambda \mapsto S_2^{(m)}(\lambda)\) is a quadratic form, so its polarisation is a bilinear prime-weighted sum to which the Chinese remainder step of Lemma 1 applies pair by pair, exactly as in Lemma 1 but without requiring the two weights to coincide.
The second is a device for controlling the size of the sieve moduli. The aggregate error bound of Lemma 1 is available only because every modulus \(q(\mathbf{d}, \mathbf{e})\) occurring in the double sum is below \(R^2 W\), hence below \(N^{1/2 - \eta}\): this is what makes Bombieri–Vinogradov applicable across the whole range. The modulus \(q(\mathbf{d}, \mathbf{e})\) is controlled by the product of the two outer truncations, one from each weight, not by the individual truncations; so if one weight is truncated more severely in its outer coordinates, the other may be truncated correspondingly less without moving the moduli at all. To exploit this one needs to split a weight into a part supported below a prescribed outer cutoff and a complementary part. The split is performed on the \(y\)-side — cut off \(y\), then transform back — rather than on the \(\lambda\)-side, because it is exactly the \(y\)-side cutoff that is transparent to the transforms and to the sup-norm \(y_{\max}\).
Definition 1 (Retained weight). Let \(\lambda\) be a finitely supported weight with \(y\)-transform \(y\), let \(m\) be an index and let \(B > 0\). The retained weight \(\lambda^{\mathrm{ret}}_{B, m}\) is the Möbius transform (Definition 1) of the cut-off family \[y^{\mathrm{ret}}_{B, m}(\mathbf{r}) \;:=\; \begin{cases} y(\mathbf{r}), & \prod_{i \ne m} r_i \le B,\\[0.3em] 0, & \text{otherwise.} \end{cases}\]
Definition 1 (Discarded weight). With \(\lambda\), \(m\) and \(B\) as in Definition 1, the discarded weight is \(\lambda^{\mathrm{disc}}_{B, m} := \lambda - \lambda^{\mathrm{ret}}_{B, m}\).
Uses: def_retained_weight
Lemma 1 (The split is transparent to the \(y\)-transform). Let \(\lambda\) have permissible support at \((R, W)\), let \(m\) be an index and let \(B > 0\). Then for every \(\mathbf{r}\), \[y_{\lambda^{\mathrm{ret}}_{B, m}}(\mathbf{r}) \;=\; \begin{cases} y(\mathbf{r}), & \prod_{i \ne m} r_i \le B,\\[0.3em] 0, & \text{otherwise,} \end{cases} \qquad\text{and}\qquad y_{\lambda^{\mathrm{disc}}_{B, m}}(\mathbf{r}) \;=\; y(\mathbf{r}) - y_{\lambda^{\mathrm{ret}}_{B, m}}(\mathbf{r}).\] In particular both pieces have permissible support at \((R, W)\) and their \(y\)-transforms have sup-norm at most \(y_{\max}\).
Uses: def_retained_weight, def_discarded_weight, lem_y_from_lambda, lem_y_permissible_support, def_adm_support, def_y_max
Proof. The \(y\)- and \(\lambda\)-transforms are mutually inverse on weights supported on tuples with squarefree coordinates (Lemma 1); applying the \(y\)-transform to \(\lambda^{\mathrm{ret}}_{B, m}\), which by Definition 1 is the \(\lambda\)-transform of \(y^{\mathrm{ret}}_{B, m}\), returns \(y^{\mathrm{ret}}_{B, m}\) on the relevant support. Linearity of the \(y\)-transform then gives the second identity. Since the cut-off family is pointwise dominated by \(y\) in absolute value, its sup-norm is at most \(y_{\max}\), and so is that of the difference after the triangle inequality is applied coordinatewise; and both transforms inherit permissible support from \(y\) by Lemma 1. ◻
Proof uses: lem_y_from_lambda, lem_lambda_from_y, lem_y_permissible_support
Lemma 1 (Additivity of the diagonal form along the split). With \(\lambda\), \(m\), \(B\) as above, \[\sum_{\mathbf{r}} \frac{Y^{(m)}(\mathbf{r})^2}{\prod_{i \ne m} g(r_i)} \;=\; \sum_{\mathbf{r}} \frac{\bigl(Y^{(m)}_{\mathrm{ret}}(\mathbf{r})\bigr)^2}{\prod_{i \ne m} g(r_i)} \;+\; \sum_{\mathbf{r}} \frac{\bigl(Y^{(m)}_{\mathrm{disc}}(\mathbf{r})\bigr)^2}{\prod_{i \ne m} g(r_i)},\] where \(Y^{(m)}_{\mathrm{ret}}\) and \(Y^{(m)}_{\mathrm{disc}}\) are the free-coordinate marginals (Definition 1) of \(\lambda^{\mathrm{ret}}_{B, m}\) and \(\lambda^{\mathrm{disc}}_{B, m}\).
Uses: def_retained_weight, def_discarded_weight, def_ym_marginal, lem_retained_y_transform, def_g_func
Proof. The marginal \(Y^{(m)}(\mathbf{r})\) sums \(y(\mathbf{r}\,[m \mapsto a])/\varphi(a)\) over the single free coordinate \(a\), and the cutoff \(\prod_{i \ne m} r_i \le B\) does not involve that coordinate. Hence by Lemma 1, for each fixed \(\mathbf{r}\) either \(Y^{(m)}_{\mathrm{ret}}(\mathbf{r}) = Y^{(m)}(\mathbf{r})\) and \(Y^{(m)}_{\mathrm{disc}}(\mathbf{r}) = 0\) (when the cutoff holds), or the reverse (when it fails). In either case \(Y^{(m)}(\mathbf{r})^2 = Y^{(m)}_{\mathrm{ret}}(\mathbf{r})^2 + Y^{(m)}_{\mathrm{disc}}(\mathbf{r})^2\) pointwise, and all three families are summable against \(1/\prod_{i \ne m} g(r_i)\), so the identity survives summation. ◻
Proof uses: lem_retained_y_transform
Lemma 1 (The modulus depends only on the product of the outer truncations). Let \(\lambda_1\) be supported on tuples with \(\prod_{i \ne m} d_i \le B_1\) and \(\lambda_2\) on tuples with \(\prod_{i \ne m} e_i \le B_2\). Then every pair \((\mathbf{d}, \mathbf{e})\) with \(\lambda_1(\mathbf{d})\lambda_2(\mathbf{e}) \ne 0\) and \(d_m = e_m = 1\) satisfies \[q(\mathbf{d}, \mathbf{e}) \;\le\; B_1\, B_2\, W.\]
Uses: def_q_lcm, def_W_trick, def_retained_weight
Proof. Since \(d_m = e_m = 1\) we have \([d_m, e_m] = 1\), so \(q(\mathbf{d}, \mathbf{e}) = W \prod_{i \ne m} [d_i, e_i] \le W \prod_{i \ne m} d_i e_i \le B_1 B_2 W\). ◻
The bound depends on \(B_1\) and \(B_2\) only through their product. For weights of permissible support one may take \(B_1 = B_2 = R\), recovering the range \(R^2 W\) of Sublemma 1; but the same range \(R^2 W\) — and hence the same appeal to Bombieri–Vinogradov — is available for any pair of outer truncations with \(B_1 B_2 \le R^2\), even when one of them exceeds \(R\). This is precisely what the split of Definitions 1 and 1 makes possible: the modulus range is governed by \(\vartheta < 1/2\) together with the split, and no further hypothesis relating the truncation exponents is needed.
Lemma 1 (Two-weight Chinese remainder bound). There is a constant \(C > 0\), depending only on \(k\), \(\mathcal{H}\) and \(m\), such that for all sufficiently large \(N\), every compatible \(v_0\) and all finitely supported weights \(\lambda_1, \lambda_2\) dominated by a common weight of permissible support at \((R, W)\), \[\begin{aligned} &\Bigl| \tfrac{1}{2}\bigl(S_2^{(m)}(\lambda_1 + \lambda_2) - S_2^{(m)}(\lambda_1) - S_2^{(m)}(\lambda_2)\bigr) \;-\; \frac{X_N}{\varphi(W)} \sideset{}{'}\sum_{\substack{\mathbf{d}, \mathbf{e}\\ d_m = e_m = 1}} \frac{\lambda_1(\mathbf{d})\,\lambda_2(\mathbf{e})}{\prod_{i=1}^k \varphi([d_i, e_i])} \Bigr|\\ &\;\le\; C \sum_{\substack{\mathbf{d}, \mathbf{e}\\ d_m = e_m = 1}} |\lambda_1(\mathbf{d})|\,|\lambda_2(\mathbf{e})|\, E\bigl(N, q(\mathbf{d}, \mathbf{e})\bigr), \end{aligned}\] where \(\sideset{}{'}\sum\) restricts to pairs for which \(W,[d_1,e_1],\dotsc,[d_k,e_k]\) are pairwise coprime and \((h_m-h_i,[d_i,e_i])=1\) for every \(i\ne m\).
Uses: lem_S2m_CRT, lem_S2m_dm_em_one, lem_S2m_expansion, lem_S2m_main_CRT, def_S2m, def_q_lcm, not_XN, not_error, def_adm_support, def_W_trick
Proof. The map \(\lambda \mapsto S_2^{(m)}(\lambda)\) is a quadratic form: by Lemma 1 it is the double sum of \(\lambda(\mathbf{d})\lambda(\mathbf{e})\,T_m(\mathbf{d}, \mathbf{e})\) with a kernel \(T_m\) not depending on \(\lambda\). Its polarisation, the left-hand quantity, is therefore \(\sum_{\mathbf{d}, \mathbf{e}} \lambda_1(\mathbf{d})\lambda_2(\mathbf{e})\, T_m(\mathbf{d}, \mathbf{e})\).
From here the proof of Lemma 1 applies verbatim, pair by pair, since it never used that the two weights coincide: Lemma 1 discards the pairs with \(d_m \ne 1\) or \(e_m \ne 1\); the pairs violating the pairwise-coprimality or shift-coprimality conditions contribute \(0\) for \(N\) large, because every prime factor of a \([d_i, e_i]\) occurring on the common dominating support exceeds \(D_0\); and Lemma 1 evaluates each retained \(T_m(\mathbf{d}, \mathbf{e})\) as \(X_N/\varphi(q(\mathbf{d}, \mathbf{e}))\) with an error at most \(C\,E(N, q(\mathbf{d}, \mathbf{e}))\). Multiplicativity of \(\varphi\) across the pairwise coprime factorisation \(q(\mathbf{d}, \mathbf{e}) = W \prod_i [d_i, e_i]\) turns the main terms into the displayed restricted sum, and the triangle inequality collects the errors. ◻
Proof uses: lem_S2m_expansion, lem_S2m_dm_em_one, lem_S2m_CRT, lem_S2m_main_CRT
From positivity to primes, and \(L^2\)-continuity
We pass from positivity of the sieve difference to a bound on prime gaps, and establish the mean-square continuity used to compare the variational supremum with the smooth-profile asymptotics.
The sieve difference \(S(\lambda) = S_2(\lambda) - \rho\, S_1(\lambda)\) counts, over the sieve window, the excess over \(\rho\) of the number of primes among \(n + h_1, \dotsc, n + h_k\). If it is positive then some \(n\) in the window has more than \(\rho\) of the shifts prime, and, this number being an integer, at least \(\lfloor \rho + 1\rfloor\) of them. Carrying this out for each large \(N\) yields infinitely many such \(n\); since the shifts \(n + h_i\) lie in an interval of length \(\operatorname{diam}(\mathcal{H})\), infinitely many intervals of that length contain \(\lfloor\rho+1\rfloor\) primes, bounding the gap between consecutive primes. Proposition 1 records this for a general tuple and a general prime count; the two bounds follow from it by choosing a tuple.
\(M_k\) is a supremum over square-integrable profiles, whereas the sieve asymptotics for \(S_1\) and \(S_2^{(m)}\) apply only to smooth profiles supported on the simplex. The comparison rests on two facts: the functionals \(I_k\) and \(J_k^{(m)}\) are continuous for the \(L^2\) topology, and smooth simplex-supported profiles are \(L^2\)-dense. A profile that nearly attains \(M_k\) may therefore be replaced by a smooth one that nearly attains it. The asymptotic for \(S\) follows, and from it the eventual positivity of \(S\).
From \(S > 0\) to primes
Lemma 1 (\(S > 0\) for all large \(N\) gives infinitely many prime-rich translates). Let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) and \(\rho > 0\). Suppose that for every sufficiently large \(N\) there exists a finitely supported \(\lambda : \mathbb{N}^k \to \mathbb{R}\) (depending on \(N\)) with \(S_2(\lambda) - \rho\, S_1(\lambda) > 0\). Then the set \[\bigl\{\, n \in \mathbb{N} \;:\; \#\{\, i : n + h_i \in \mathbb{P}\,\} \;\ge\; \lfloor \rho + 1\rfloor \,\bigr\}\] is infinite.
Proof. Write \(\mathcal{A}\) for the displayed set. It suffices to show that \(\mathcal{A}\) is unbounded above, since an unbounded subset of \(\mathbb{N}\) is infinite.
Let \(M \in \mathbb{N}\) be arbitrary. Choose \(N \ge M\) large enough that the hypothesis supplies a finitely supported \(\lambda\) with \(S_2(\lambda) - \rho\, S_1(\lambda) > 0\) for that \(N\). Lemma 1 then produces an integer \(n\) with \(N < n \le 2N\) (and \(n \equiv v_0 \pmod W\), which we discard) such that at least \(\lfloor \rho + 1\rfloor\) of the \(n + h_i\) are prime. Thus \(n \in \mathcal{A}\) and \(n > N \ge M\). As \(M\) was arbitrary, \(\mathcal{A}\) is unbounded. ◻
Proof uses: lem_S_positive_implies_primes
Definition 1 (Prime-rich interval). For \(n, L, r \in \mathbb{N}\), say that the closed interval \([n, n + L]\) is \(r\)-prime-rich if it contains at least \(r\) distinct primes, that is, if \[\#\bigl(\, [n, n+L] \cap \mathbb{P}\,\bigr) \;\ge\; r .\]
Lemma 1 (Prime-rich translates give prime-rich intervals). Let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) with \(k \ge 1\) and \(h_1 < \dotsb < h_k\), and let \(r \in \mathbb{N}\). If the set \(\{\, n : \#\{i : n + h_i \in \mathbb{P}\} \ge r \,\}\) is infinite, then for arbitrarily large \(n'\) the interval \([n', n' + \operatorname{diam}(\mathcal{H})]\) is \(r\)-prime-rich in the sense of Definition 1.
Proof. Write \(\mathcal{A}\) for the hypothesised infinite set and \(D := \operatorname{diam}(\mathcal{H}) = h_k - h_1\). Let \(M\) be arbitrary and pick \(n \in \mathcal{A}\) with \(n > M\); put \(n' := n + h_1 \ge n > M\).
Consider the map \(i \mapsto n + h_i\) from \(\{i : n + h_i \in \mathbb{P}\}\) to \([n', n' + D] \cap \mathbb{P}\). It is well defined: if \(n + h_i\) is prime then it is a prime, and \(h_1 \le h_i \le h_k\) gives \(n' = n + h_1 \le n + h_i \le n + h_k = n' + D\). It is injective because the \(h_i\) are distinct. Hence \[\#\bigl([n', n' + D] \cap \mathbb{P}\bigr) \;\ge\; \#\{i : n + h_i \in \mathbb{P}\} \;\ge\; r ,\] so \([n', n' + D]\) is \(r\)-prime-rich. Since \(M\) was arbitrary and \(n' > M\), such \(n'\) occur arbitrarily late. ◻
Lemma 1 (Prime-rich intervals imply small prime gaps). Let \(r \ge 1\) and \(L \ge 0\) be integers. If the interval \([n, n + L]\) is \(r\)-prime-rich for arbitrarily large \(n\), then \[p_{m + r - 1} - p_m \;\le\; L\] for infinitely many \(m\).
Proof. For \(n \ge 1\) let \(\mu(n) := \min\{\, j \ge 1 : p_j \ge n \,\}\) be the index of the least prime that is at least \(n\); this is well defined because there are infinitely many primes. Since exactly \(\pi(n-1)\) primes are smaller than \(n\), we have \(\mu(n) = \pi(n-1) + 1\), so \(\mu\) is non-decreasing and \(\mu(n) \to \infty\) as \(n \to \infty\).
Let \(M\) be arbitrary; we produce \(m \ge M\) with \(p_{m + r - 1} - p_m \le L\). Choose \(n_0\) with \(\mu(n_0) \ge M\), and then, by hypothesis, some \(n \ge n_0\) for which \([n, n+L]\) is \(r\)-prime-rich. Put \(m := \mu(n)\), so that \(m \ge \mu(n_0) \ge M\).
By construction \(p_m \ge n\), and every prime in \([n, n+L]\) has index at least \(m\). Since the primes in \([n, n+L]\) are consecutive in the enumeration, they are \(p_m, p_{m+1}, \dotsc, p_{\pi(n+L)}\), and there are at least \(r\) of them, so \(\pi(n + L) \ge m + r - 1\) and therefore \(p_{m + r - 1} \le n + L\). Subtracting, \[p_{m + r - 1} - p_m \;\le\; (n + L) - n \;=\; L .\] As \(M\) was arbitrary, this holds for infinitely many \(m\). ◻
Proposition 1 (Bounded gaps from prime-rich translates). Let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) with \(k \ge 1\) and \(h_1 < \dotsb < h_k\), and let \(r \ge 1\) be an integer. If the set \[\bigl\{\, n \in \mathbb{N} \;:\; \#\{\, i : n + h_i \in \mathbb{P}\,\} \;\ge\; r \,\bigr\}\] is infinite, then \(p_{m + r - 1} - p_m \le \operatorname{diam}(\mathcal{H})\) for infinitely many \(m\).
Uses: def_diam
Proof. Apply Lemma 1 to obtain, for arbitrarily large \(n'\), an \(r\)-prime-rich interval \([n', n' + \operatorname{diam}(\mathcal{H})]\); then apply Lemma 1 with \(L = \operatorname{diam}(\mathcal{H})\). ◻
\(L^2\)-continuity of the functionals
Throughout this subsection \(k \ge 1\) and, for \(s > 0\), \(\mathcal{R}_k(s) := \{ t \in [0,\infty)^k : t_1 + \dotsb + t_k \le s\}\) denotes the simplex of scale \(s\), so that \(\mathcal{R}_k = \mathcal{R}_k(1)\) and the enlarged simplex \(\mathcal{T}_\varepsilon\) is \(\mathcal{R}_k(1 + \varepsilon)\). Every \(\mathcal{R}_k(s)\) is compact, being closed and contained in the ball of radius \(s\) about the origin. Functions in \(L^2(\mathcal{R}_k(s))\) are always regarded as defined on all of \(\mathbb{R}^k\) by extension by zero.
Definition 1 (Partial-integration operator \(T_m\)). Let \(s > 0\) and \(m \in \{1, \dotsc, k\}\). The \(m\)-th partial-integration operator is the linear map \[T_m : L^2(\mathcal{R}_k(s)) \longrightarrow L^2(\mathbb{R}^{k-1}), \qquad (T_m f)(t_1, \dotsc, \widehat{t_m}, \dotsc, t_k) \;:=\; \int_{\mathbb{R}} f(t_1, \dotsc, t_k)\, dt_m ,\] the hat denoting omission of the \(m\)-th coordinate.
Uses: def_simplex
Smooth approximation on the simplex
Definition 1 (The main scale). For \(k \in \mathbb{N}\) and \(N \in \mathbb{N}\) set \[\mathfrak{M}_k(N) \;:=\; \frac{\varphi(W)^k\, N\, (\log R)^k}{W^{k+1}} ,\] the common scale of the first and second sieve moments.
Uses: def_W_trick, def_R
Proposition 1 (Simultaneous asymptotics for the sieve moments). Let \(k \ge 2\), let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) be admissible, and let \(F \in C^\infty(\mathbb{R}^k)\) have \(\operatorname{supp} F \subset \mathcal{R}_k\). Let \(0 < \vartheta < 1/2\) and \(0 < \delta < \vartheta/2\), and suppose the primes have level of distribution \(\vartheta\). Write \(\lambda_F\) for the sieve weight attached to \(F\) at truncation \(R = N^{\vartheta/2 - \delta}\). Then there are functions \(e_1, e_2 : \mathbb{N} \to \mathbb{R}\) with \(e_1(N) \to 0\) and \(e_2(N) \to 0\), and a threshold \(N_0\), such that for every \(N \ge N_0\) and every residue \(v_0\) modulo \(W\) with \((v_0 + h_i, W) = 1\) for all \(i\), the two estimates \[\bigl| S_1(\lambda_F) - \mathfrak{M}_k(N)\, I_k(F) \bigr| \;\le\; e_1(N)\, \mathfrak{M}_k(N)\] and \[\Bigl| S_2(\lambda_F) - \frac{\varphi(W)^k\, N\, (\log R)^{k+1}}{W^{k+1}\log N} \sum_{m=1}^{k} J_k^{(m)}(F) \Bigr| \;\le\; e_2(N)\, \mathfrak{M}_k(N)\] hold simultaneously.
Uses: def_admissible, def_simplex, def_level_of_distribution, def_lambda_from_F, def_R, def_S1, def_S2, def_I_k, def_J_k, def_main_scale, def_W_trick
Proof. The two estimates are the smooth-profile asymptotics for the two moments, made uniform in \(N\) and in \(v_0\).
For the first, Lemma 1 supplies a constant \(C_1 > 0\) and a threshold \(N_1\) such that for all \(N \ge N_1\) and all admissible residues \(v_0\), \[\Bigl| S_1(\lambda_F) - \frac{\varphi(W)^k N (\log R)^k}{W^{k+1}} \int_{\mathcal{R}_k} F^2 \Bigr| \;\le\; \frac{C_1 F_{\max}^2\, \varphi(W)^k N (\log R)^k}{W^{k+1} D_0} .\] The main term is \(\mathfrak{M}_k(N)\, I_k(F)\) by Definitions 1 and 1, and the error is \(e_1(N)\,\mathfrak{M}_k(N)\) with \(e_1(N) := C_1 F_{\max}^2 / D_0\). Since \(D_0 = \log\log\log N \to \infty\), we have \(e_1(N) \to 0\).
For the second, Lemma 1 supplies, for each \(m \in \{1, \dotsc, k\}\), a constant \(C^{(m)} > 0\) and a threshold \(N^{(m)}\) such that for all \(N \ge N^{(m)}\) and all admissible \(v_0\), \[\Bigl| S_2^{(m)}(\lambda_F) - \frac{\varphi(W)^k N (\log R)^{k+1}}{W^{k+1}\log N}\, J_k^{(m)}(F) \Bigr| \;\le\; \frac{C^{(m)} F_{\max}^2\, \varphi(W)^k N (\log R)^k}{W^{k+1} D_0} .\] Summing over the finitely many \(m\) and using the decomposition \(S_2(\lambda) = \sum_{m=1}^k S_2^{(m)}(\lambda)\) of Lemma 1, the triangle inequality gives the second estimate with \(e_2(N) := \bigl(\sum_m C^{(m)}\bigr) F_{\max}^2 / D_0\), which again tends to \(0\).
Finally take \(N_0\) to be the largest of \(N_1\) and the finitely many \(N^{(m)}\) (and large enough that \(D_0 \ge 2\), so that the error scales are meaningful). Both estimates then hold for all \(N \ge N_0\), simultaneously and uniformly in \(v_0\). ◻
Proof uses: lem_S1_smooth, lem_S2m_smooth, lem_S2_decomp, lem_Ik_quadratic, def_D_0, def_S2m
Lemma 1 (Asymptotic for the sieve difference). Under the hypotheses of Proposition 1 and for any \(\rho > 0\), there is a function \(e : \mathbb{N} \to \mathbb{R}\) with \(e(N) \to 0\) and a threshold \(N_0\) such that for every \(N \ge N_0\) and every residue \(v_0\) modulo \(W\) with \((v_0 + h_i, W) = 1\) for all \(i\), \[\Bigl| S(\lambda_F) - \mathfrak{M}_k(N)\Bigl( \bigl(\tfrac{\vartheta}{2} - \delta\bigr) \sum_{m=1}^{k} J_k^{(m)}(F) - \rho\, I_k(F) \Bigr) \Bigr| \;\le\; e(N)\, \mathfrak{M}_k(N) ,\] where \(S(\lambda_F) = S_2(\lambda_F) - \rho\, S_1(\lambda_F)\).
Uses: prop_main_prop, def_main_scale, def_I_k, def_J_k, def_S1, def_S2, def_R
Proof. The point is that the two main terms of Proposition 1 share the scale \(\mathfrak{M}_k(N)\). Indeed \(\log R = (\vartheta/2 - \delta)\log N\) by the definition \(R = N^{\vartheta/2 - \delta}\), so \[\frac{\varphi(W)^k N (\log R)^{k+1}}{W^{k+1}\log N} \;=\; \frac{\varphi(W)^k N (\log R)^{k}}{W^{k+1}} \cdot \frac{\log R}{\log N} \;=\; \mathfrak{M}_k(N)\,\Bigl(\frac{\vartheta}{2} - \delta\Bigr).\]
Let \(e_1, e_2, N_0\) be as in Proposition 1 and put \(e := e_2 + \rho\, e_1\), which tends to \(0\). Writing \(P := \mathfrak{M}_k(N)(\vartheta/2 - \delta)\), \(\mathcal{J} := \sum_m J_k^{(m)}(F)\) and \(\mathcal{I} := I_k(F)\), we have the algebraic identity \[\bigl(S_2(\lambda_F) - \rho\, S_1(\lambda_F)\bigr) - \mathfrak{M}_k(N)\bigl((\vartheta/2 - \delta)\mathcal{J} - \rho\,\mathcal{I}\bigr) \;=\; \bigl(S_2(\lambda_F) - P\,\mathcal{J}\bigr) - \rho\,\bigl(S_1(\lambda_F) - \mathfrak{M}_k(N)\,\mathcal{I}\bigr) .\] The triangle inequality and the two estimates of Proposition 1 bound the absolute value of the right-hand side by \[e_2(N)\,\mathfrak{M}_k(N) + \rho\, e_1(N)\,\mathfrak{M}_k(N) \;=\; e(N)\, \mathfrak{M}_k(N),\] as required. ◻
Lemma 1 (Eventual positivity of the sieve difference). Assume the hypotheses of Lemma 1, and assume in addition the strict inequality \[\rho\, I_k(F) \;<\; \Bigl(\frac{\vartheta}{2} - \delta\Bigr) \sum_{m=1}^{k} J_k^{(m)}(F) .\] Then there is a threshold \(N_0\) such that for every \(N \ge N_0\) and every residue \(v_0\) modulo \(W\) with \((v_0 + h_i, W) = 1\) for all \(i\), \[S_2(\lambda_F) - \rho\, S_1(\lambda_F) \;>\; 0 .\]
Uses: lem_S_asymptotic, def_I_k, def_J_k, def_S1, def_S2, def_main_scale
Proof. Set \(c := (\vartheta/2 - \delta)\sum_m J_k^{(m)}(F) - \rho\, I_k(F)\), which is strictly positive by hypothesis. Let \(e\) and \(N_1\) be as in Lemma 1. Since \(e(N) \to 0\), there is \(N_2\) with \(e(N) < c\) for all \(N \ge N_2\).
The scale \(\mathfrak{M}_k(N) = \varphi(W)^k N (\log R)^k / W^{k+1}\) is strictly positive once \(N\) is large enough that \(\log R > 0\), which holds as soon as \(N > 1\), because \(\vartheta/2 - \delta > 0\). Let \(N_0\) be the largest of \(N_1\), \(N_2\) and this threshold.
For \(N \ge N_0\) and any admissible \(v_0\), the lower half of the estimate of Lemma 1 gives \[S_2(\lambda_F) - \rho\, S_1(\lambda_F) \;\ge\; \mathfrak{M}_k(N)\, c - e(N)\,\mathfrak{M}_k(N) \;=\; \mathfrak{M}_k(N)\,\bigl(c - e(N)\bigr) \;>\; 0 ,\] since \(\mathfrak{M}_k(N) > 0\) and \(c - e(N) > 0\). ◻
Proposition 1 (The criterion on the standard simplex). Let \(k \ge 2\), let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) be admissible, let \(\rho > 0\), and let \(\vartheta \in (0,1)\) be a level of distribution for the primes. Suppose that \[M_k \;>\; \frac{2\rho}{\vartheta} .\] Then the set \(\{\, n \in \mathbb{N} : \#\{ i : n + h_i \in \mathbb{P}\} \ge \lfloor \rho \rfloor + 1 \,\}\) is infinite, and consequently \(p_{m + \lfloor \rho \rfloor} - p_m \le \operatorname{diam}(\mathcal{H})\) for infinitely many \(m\).
Proof. Apply Lemma 1 at \(\varepsilon = 0\). Since \(M_k > 2\rho/\vartheta\), it produces a smooth \(F\) supported in the closed simplex \(\mathcal{R}_k\), together with a \(\delta \in (0, \vartheta/2)\), such that \[\rho\, I_k(F) \;<\; \Bigl(\frac{\vartheta}{2} - \delta\Bigr) \sum_{m=1}^{k} J_k^{(m)}(F) .\] Because \(F\) is supported in the standard simplex, no enlargement of the sieve support is required, and the tensor coefficients derived from \(F\) as in Definition 1 are admissible at truncation \(R = N^{\vartheta/2 - \delta}\).
The strict inequality above is precisely the additional hypothesis of Lemma 1; its remaining hypotheses are those of Lemma 1, which hold for this \(F\) and this \(\delta\) at the level \(\vartheta\). Lemma 1 therefore yields an \(N_0\) such that \(S_2(\lambda) - \rho\, S_1(\lambda) > 0\) for every \(N \ge N_0\), where \(\lambda\) is the coefficient system attached to \(F\) at scale \(N\).
Each such \(\lambda\) is finitely supported, so Lemma 1 applies and shows that \(\{\, n : \#\{ i : n + h_i \in \mathbb{P}\} \ge \lfloor \rho + 1 \rfloor \,\}\) is infinite. Since \(\lfloor \rho + 1 \rfloor = \lfloor \rho \rfloor + 1\) for every real \(\rho\), this gives the first assertion. The second follows by Proposition 1 applied with \(r = \lfloor \rho \rfloor + 1\).
The hypothesis \(M_k > 2\rho/\vartheta\) permits every \(\rho < \vartheta M_k/2\), and the count \(\lfloor \rho \rfloor + 1\) is largest when \(\rho\) is taken just below that threshold: for \(\rho = \vartheta M_k/2 - \eta\) with \(\eta > 0\) small enough one has \(\lfloor \rho \rfloor + 1 = \lceil \vartheta M_k/2 \rceil = r_k(\vartheta)\), the prime count of Definition 1. Thus \(r_k(\vartheta)\) is exactly the number of primes this proposition delivers at level \(\vartheta\). ◻
Instantiation I: the standard simplex and the bound \(H_1 \le 600\)
We fix the data for the first case. The domain is the standard simplex \(\mathcal{R}_k\); the dimension is \(k = 105\); the profile is an explicit symmetric polynomial supported on \(\mathcal{R}_{105}\); and the admissible tuple is an explicit \(105\)-tuple of diameter \(600\). The variational criterion requires \(M_{105} > 4\), and the profile certifies it. The only analytic input is the Bombieri–Vinogradov theorem (Theorem 1); no hypothesis on the level of distribution beyond \(\vartheta < 1/2\) is assumed.
The tuple \(\mathcal{H}_{105}\)
Definition 1 (The \(105\)-tuple). Let \[\begin{aligned} \mathcal{H}_{105} \;:=\; \{\, &0,\, 10,\, 12,\, 24,\, 28,\, 30,\, 34,\, 42,\, 48,\, 52,\, 54,\, 64,\, 70,\\ &72,\, 78,\, 82,\, 90,\, 94,\, 100,\, 112,\, 114,\, 118,\, 120,\, 124,\, 132,\, 138,\\ &148,\, 154,\, 168,\, 174,\, 178,\, 180,\, 184,\, 190,\, 192,\, 202,\, 204,\, 208,\, 220,\\ &222,\, 232,\, 234,\, 250,\, 252,\, 258,\, 262,\, 264,\, 268,\, 280,\, 288,\, 294,\, 300,\\ &310,\, 322,\, 324,\, 328,\, 330,\, 334,\, 342,\, 352,\, 358,\, 360,\, 364,\, 372,\, 378,\\ &384,\, 390,\, 394,\, 400,\, 402,\, 408,\, 412,\, 418,\, 420,\, 430,\, 432,\, 442,\, 444,\\ &450,\, 454,\, 462,\, 468,\, 472,\, 478,\, 484,\, 490,\, 492,\, 498,\, 504,\, 510,\, 528,\\ &532,\, 534,\, 538,\, 544,\, 558,\, 562,\, 570,\, 574,\, 580,\, 582,\, 588,\, 594,\, 598,\\ &600\,\}, \end{aligned}\] a set of \(105\) non-negative integers listed in increasing order, with least element \(0\) and greatest element \(600\). It is the narrowest admissible \(105\)-tuple recorded in the tables of Engelsma and Sutherland.
Lemma 1 (Cardinality of \(\mathcal{H}_{105}\)). \(\#\mathcal{H}_{105} = 105\).
Uses: def_h105
Proof. The list in Definition 1 is strictly increasing, hence has pairwise distinct entries; counting them gives \(105\). ◻
Lemma 1 (Admissibility of \(\mathcal{H}_{105}\)). \(\mathcal{H}_{105}\) is admissible in the sense of Definition 1.
Uses: def_h105, def_admissible
Proof. By Lemma 1 it suffices to exhibit, for each prime \(p \le \#\mathcal{H}_{105} = 105\), a residue class modulo \(p\) omitted by \(\mathcal{H}_{105}\): for a prime \(p > 105\) the reduction map \(\mathcal{H}_{105} \to \mathbb{Z}/p\mathbb{Z}\) has image of cardinality at most \(\#\mathcal{H}_{105} = 105 < p\) by Lemma 1, so by Lemma 1 the omission is automatic. There are exactly twenty-seven primes \(p \le 105\), namely \[2,\,3,\,5,\,7,\,11,\,13,\,17,\,19,\,23,\,29,\,31,\,37,\,41,\,43,\,47,\,53,\,59,\,61,\,67,\,71,\,73,\,79,\,83,\,89,\,97,\,101,\,103,\] and for each the set \(\{h \bmod p : h \in \mathcal{H}_{105}\}\) is an explicit finite subset of \(\mathbb{Z}/p\mathbb{Z}\) which may be computed by direct enumeration of the \(105\) listed integers; in every case the enumeration omits a class. (For instance every element of \(\mathcal{H}_{105}\) is even, so the class \(1 \bmod 2\) is omitted; the classes \(2 \bmod 3\), \(1 \bmod 5\) and \(4 \bmod 7\) are likewise omitted.) Since the check is a finite computation on an explicit finite set, it is decidable. ◻
Proof uses: lem_adm_small_primes, lem_admissible_cardinality_char, lem_h105_card
Lemma 1 (Diameter of \(\mathcal{H}_{105}\)). \(\operatorname{diam}(\mathcal{H}_{105}) = 600\).
Proof. By Lemma 1, \(\operatorname{diam}(\mathcal{H}_{105}) = \max \mathcal{H}_{105} - \min\mathcal{H}_{105}\). The list of Definition 1 is increasing, so its minimum is its first entry \(0\) and its maximum is its last entry \(600\); the difference is \(600\). ◻
Proof uses: lem_diam_max_min
The variational datum at \(k = 105\)
The profile lies in the space spanned by the symmetric monomials \((1 - P_1)^b P_2^{c}\) of the two power sums \(P_1 = t_1 + \dotsb + t_{105}\) and \(P_2 = t_1^2 + \dotsb + t_{105}^2\), restricted to total weight \(b + 2c \le 11\); this is the basis \(\mathcal{B}_{11}\) of Definition 1. A profile in this span is determined by its coefficient vector, and both \(I_{105}\) and \(J^{(m)}_{105}\) restrict to explicit rational quadratic forms in that vector, so a single strict inequality between two rational numbers certifies the required bound.
Definition 1 (The certificate profile at \(k = 105\)). Let \((b^*_i, c^*_i;\, a^*_i)_{1 \le i \le 32}\) be the following list of triples, in which \(b^*_i, c^*_i \in \mathbb{Z}_{\ge 0}\) satisfy \(b^*_i + 2 c^*_i \le 11\) and \(a^*_i \in \mathbb{Z}\): \[\begin{aligned} &(1,0;\,{-}330),\ (1,1;\,78029),\ (4,0;\,5743874),\\ &(1,2;\,{-}7297879),\ (0,3;\,16574),\ (2,2;\,{-}12057779),\\ &(4,1;\,{-}1391698982),\ (6,0;\,{-}1110129088),\ (1,3;\,340699862),\\ &(3,2;\,896499841),\ (5,1;\,6664560083),\ (7,0;\,3420088505),\\ &(0,4;\,{-}1685613),\ (2,3;\,1231124500),\ (4,2;\,98419064701),\\ &(6,1;\,126065419782),\ (1,4;\,{-}8050576592),\ (3,3;\,{-}90260561416),\\ &(5,2;\,{-}636913875947),\ (9,0;\,305538813668),\ (0,5;\,42387884),\\ &(2,4;\,{-}33232242480),\ (4,3;\,{-}2440220187876),\ (6,2;\,{-}7132620815184),\\ &(8,1;\,{-}6762976569892),\ (10,0;\,{-}1975749959806),\ (1,5;\,78797053343),\\ &(3,4;\,2504524184641),\ (5,3;\,24468760733568),\ (7,2;\,35184372088832),\\ &(9,1;\,18268089505780),\ (11,0;\,3188339316861). \end{aligned}\] The certificate polynomial at \(k = 105\) is the symmetric polynomial \[P^*(t_1,\dotsc,t_{105}) \;:=\; \sum_{i=1}^{32} a^*_i \,\bigl(1 - P_1\bigr)^{b^*_i} \,P_2^{\,c^*_i}, \qquad P_j \;=\; \sum_{\ell = 1}^{105} t_\ell^{\,j}.\]
Uses: def_basis_Bd
Definition 1 (The \(I\)-value of the certificate). With the data of Definition 1, set \[A^* \;:=\; \sum_{i=1}^{32}\sum_{j=1}^{32} a^*_i\,a^*_j\; \frac{(b^*_i + b^*_j)! \; G_{c^*_i + c^*_j,\,2}(105)} {\bigl(105 + b^*_i + b^*_j + 2c^*_i + 2c^*_j\bigr)!} \;\in\; \mathbb{Q},\] where \(G_{b,2}\) is the polynomial of Definition 1.
Definition 1 (The \(J\)-value of the certificate). With the data of Definition 1, set \[B^* \;:=\; \sum_{i=1}^{32}\sum_{j=1}^{32} a^*_i\,a^*_j\; \frac{\displaystyle\sum_{c' = 0}^{c^*_i}\ \sum_{c'' = 0}^{c^*_j} \binom{c^*_i}{c'}\binom{c^*_j}{c''}\, \gamma\bigl(b^*_i, b^*_j, c^*_i, c^*_j, c', c''\bigr)\, G_{c' + c'',\,2}(104)} {\bigl(105 + b^*_i + b^*_j + 2c^*_i + 2c^*_j + 1\bigr)!} \;\in\; \mathbb{Q},\] where \(G_{b,2}\) is the polynomial of Definition 1 and, for non-negative integers \(b_1, b_2, c_1, c_2\) and \(c' \le c_1\), \(c'' \le c_2\), \[\gamma(b_1, b_2, c_1, c_2, c', c'') \;:=\; \frac{b_1!\; b_2!\; (2c_1 - 2c')!\; (2c_2 - 2c'')!\; \bigl(b_1 + b_2 + 2c_1 + 2c_2 - 2c' - 2c'' + 2\bigr)!} {\bigl(b_1 + 2c_1 - 2c' + 1\bigr)!\;\bigl(b_2 + 2c_2 - 2c'' + 1\bigr)!}.\]
Lemma 1 (The certificate profile realises \(A^*\)). Let \(F_{P^*}\) be the extension of \(P^*\) by zero outside \(\mathcal{R}_{105}\) (Definition 1). Then \[I_{105}\bigl(F_{P^*}\bigr) \;=\; A^*.\]
Uses: def_p105, def_a105, def_polynomial_F, def_I_k
Proof. By Definition 1, \(I_{105}(F_{P^*}) = \int_{\mathcal{R}_{105}} (P^*)^2\). Expanding the square of the finite sum of Definition 1 and using \((1 - P_1)^{b^*_i} P_2^{c^*_i} \cdot (1 - P_1)^{b^*_j} P_2^{c^*_j} = (1 - P_1)^{b^*_i + b^*_j} P_2^{c^*_i + c^*_j}\), we get \[I_{105}\bigl(F_{P^*}\bigr) \;=\; \sum_{i,j} a^*_i a^*_j \int_{\mathcal{R}_{105}} (1 - P_1)^{b^*_i + b^*_j} P_2^{\,c^*_i + c^*_j}.\] Lemma 1 evaluates each of these simplex integrals in closed form, \[\int_{\mathcal{R}_k} (1 - P_1)^{b} P_2^{\,c} \;=\; \frac{b!\; G_{c,2}(k)}{(k + b + 2c)!},\] and substituting \(k = 105\), \(b = b^*_i + b^*_j\), \(c = c^*_i + c^*_j\) reproduces exactly the double sum of Definition 1. ◻
Proof uses: lem_Ik_quadratic
Lemma 1 (The certificate profile realises \(B^*\) in every coordinate). Let \(F_{P^*}\) be the extension of \(P^*\) by zero outside \(\mathcal{R}_{105}\) (Definition 1). Then for every \(m \in \{1, \dotsc, 105\}\), \[J^{(m)}_{105}\bigl(F_{P^*}\bigr) \;=\; B^*.\]
Uses: def_p105, def_b105, def_polynomial_F, def_J_k
Proof. Fix \(m\). Since \(F_{P^*}\) vanishes off \(\mathcal{R}_{105}\), for \((t_i)_{i \ne m} \in \mathcal{R}_{105}^{(m)}\) the inner integral \(\int_0^\infty F_{P^*}\,dt_m\) is the integral of \(P^*\) over the segment \(0 \le t_m \le 1 - \sum_{i \ne m} t_i\). The integrand is linear in the coefficient vector, so squaring and integrating over \(\mathcal{R}_{105}^{(m)}\) exhibits \(J^{(m)}_{105}(F_{P^*})\) as the quadratic form \(\sum_{i,j} a^*_i a^*_j \Lambda_m(b^*_i, b^*_j, c^*_i, c^*_j)\), where \(\Lambda_m\) is the pairing of the \(m\)-th marginals of two basis monomials. Lemma 1 evaluates \(\Lambda_m\) in closed form: the inner one-dimensional integral of \((1 - P_1)^{b} P_2^{c}\) in the variable \(t_m\) is computed by expanding \(P_2 = t_m^2 + \sum_{i \ne m} t_i^2\) binomially and integrating each resulting Beta-type integral, and the remaining \((105-1)\)-dimensional integral over \(\mathcal{R}_{105}^{(m)}\) is again a simplex integral of the shape appearing in Lemma 1, now in dimension \(104\). The value obtained is \[\Lambda_m(b_1, b_2, c_1, c_2) \;=\; \frac{\sum_{c' \le c_1} \sum_{c'' \le c_2} \binom{c_1}{c'}\binom{c_2}{c''}\, \gamma(b_1, b_2, c_1, c_2, c', c'')\, G_{c' + c'',2}(104)} {(105 + b_1 + b_2 + 2c_1 + 2c_2 + 1)!},\] with \(\gamma\) as in Definition 1. This expression does not involve \(m\): the symmetry of \(P^*\) in \(t_1, \dotsc, t_{105}\) means the marginal in any one coordinate is obtained from the marginal in any other by a permutation of the remaining variables, which preserves \(\mathcal{R}_{105}^{(m)}\) and the Lebesgue measure. Summing over \(i, j\) gives \(J^{(m)}_{105}(F_{P^*}) = B^*\) for every \(m\). ◻
Proof uses: lem_quad_forms, lem_Ik_quadratic
Lemma 1 (The certificate inequality). \(4 A^* < 105\, B^*\).
Proof. Both \(A^*\) and \(B^*\) are explicit rational numbers: each is a finite double sum of \(32^2\) terms whose entries are the integers \(a^*_i\) and quotients of factorials and values \(G_{c,2}(105)\), \(G_{c,2}(104)\), all of which are determined by Definitions 1 and 1 and computable by exact integer arithmetic. Every denominator occurring in \(A^*\) divides \((105 + L + 2M)!\), and every denominator occurring in \(B^*\) divides the same factorial, where \(L := 2\max_i b^*_i + 1 = 23\) and \(M := 2\max_i c^*_i + 1 = 11\); multiplying through by this single positive common denominator turns the asserted inequality into a comparison of two integers, \[4 \cdot \bigl[(105 + L + 2M)!\, A^*\bigr] \;<\; 105 \cdot \bigl[(105 + L + 2M)!\, B^*\bigr],\] both sides of which are integers built from the \(a^*_i\) by additions and multiplications. That integer comparison holds, and dividing back by the positive common denominator gives the claim. ◻
Proposition 1 (The spectral bound at \(k = 105\)). \(M_{105} > 4\).
Uses: def_M_k
Proof. Write \(F := F_{P^*}\) for the profile of Lemma 1. It is a polynomial multiplied by the indicator of the compact set \(\mathcal{R}_{105}\), hence bounded and measurable, hence square-integrable on \(\mathcal{R}_{105}\); it is symmetric in \(t_1,\dotsc,t_{105}\) and supported in \(\mathcal{R}_{105}\). Thus \(F\) is a competitor in the supremum defining \(M_{105}\) (Definition 1), and consequently \[M_{105} \;\ge\; \frac{\sum_{m=1}^{105} J^{(m)}_{105}(F)}{I_{105}(F)}\] as soon as \(I_{105}(F) \ne 0\).
We first check \(I_{105}(F) > 0\). By Lemma 1, \(I_{105}(F) = A^*\), and \(A^* = \int_{\mathcal{R}_{105}} (P^*)^2 \ge 0\). Suppose \(A^* = 0\). Then \(P^*\) vanishes almost everywhere on \(\mathcal{R}_{105}\), so \(F = 0\) almost everywhere, so every marginal \(\int_0^\infty F \, dt_m\) vanishes almost everywhere and \(J^{(m)}_{105}(F) = 0\) for each \(m\); by Lemma 1 this forces \(B^* = 0\). But then Lemma 1 reads \(0 < 0\), a contradiction. Hence \(A^* > 0\).
By Lemma 1 the \(105\) marginal functionals all take the value \(B^*\), so \(\sum_{m=1}^{105} J^{(m)}_{105}(F) = 105\,B^*\). Combining with Lemma 1 and \(A^* > 0\), \[M_{105} \;\ge\; \frac{105\,B^*}{A^*} \;>\; \frac{4A^*}{A^*} \;=\; 4. \qedhere\] ◻
Proof uses: lem_i105_value, lem_j105_value, lem_cert105, def_p105, def_a105, def_b105
The bound \(H_1 \le 600\)
The criterion of Proposition 1 states that for an admissible \(k\)-tuple, a level of distribution \(\vartheta\) and a parameter \(\rho > 0\) with \(M_k > 2\rho/\vartheta\), there are infinitely many \(n\) for which at least \(\lfloor \rho \rfloor + 1\) of the shifts \(n + h_i\) are prime. To extract two primes it suffices to take \(\rho = 1\), which requires \[M_k \;>\; \frac{2}{\vartheta}.\] Bombieri–Vinogradov gives every \(\vartheta < 1/2\) and nothing more, so the threshold is \(\displaystyle\inf_{\vartheta < 1/2} 2/\vartheta = 4\): the criterion yields two primes precisely when \(M_k > 4\), and yields nothing when \(M_k \le 4\). This is the exact threshold certified by Proposition 1. The next lemma converts the strict inequality \(M > 4\) into an admissible choice of \(\vartheta\), with no limiting argument in \(\vartheta\).
Lemma 1 (A level of distribution matched to a spectral bound). Let \(M \in \mathbb{R}\) with \(M > 4\). Then there exists \(\vartheta\) with \(0 < \vartheta < 1/2\) and \(\lceil \vartheta M / 2 \rceil \ge 2\).
Proof. Take \(\vartheta := \tfrac14 + \tfrac1M\). Since \(M > 4 > 0\) we have \(\vartheta > 0\), and \(1/M < 1/4\) gives \(\vartheta < \tfrac14 + \tfrac14 = \tfrac12\). Moreover \[\frac{\vartheta M}{2} \;=\; \frac{1}{2}\Bigl(\frac{M}{4} + 1\Bigr) \;=\; \frac{M}{8} + \frac12 \;>\; \frac48 + \frac12 \;=\; 1,\] using \(M > 4\). A real number strictly greater than \(1\) has ceiling at least \(2\), so \(\lceil \vartheta M / 2\rceil \ge 2\). ◻
Theorem 4 (Bounded gaps: the bound \(600\)). \[\liminf_{n \to \infty}\bigl(p_{n+1} - p_n\bigr) \;\le\; 600 .\]
Proof. By Proposition 1 we have \(M_{105} > 4\), so Lemma 1 applied with \(M = M_{105}\) produces \(\vartheta \in (0, 1/2)\) with \(r_{105} = \lceil \vartheta M_{105} / 2 \rceil \ge 2\). For this \(\vartheta\) the Bombieri–Vinogradov theorem (Theorem 1) supplies the level of distribution \(\vartheta\) required by the criterion.
Write \(\mathcal{H}_{105} = \{h_1 < h_2 < \dotsb < h_{105}\}\); by Lemma 1 this is a tuple of length \(k = 105 \ge 2\), and by Lemma 1 it is admissible. Since \(M_{105} > 4\) and \(\vartheta\) was chosen with \(2/\vartheta < M_{105}\), the hypothesis \(M_{105} > 2\rho/\vartheta\) of Proposition 1 holds at \(\rho = 1\). Applying that proposition at \(k = 105\) with this tuple, this \(\vartheta\) and \(\rho = 1\) yields \(p_{m+1} - p_m \le \operatorname{diam}(\mathcal{H}_{105})\) for infinitely many \(m\).
By Lemma 1 that diameter is \(600\), so \[\liminf_{m \to \infty}\bigl(p_{m+1} - p_m\bigr) \;\le\; 600. \qedhere\] ◻
Proof uses: prop_M105, lem_theta_from_M, lem_h105_card, lem_h105_adm, lem_h105_diam, prop_criterion, thm_BV
The enlarged variational data, sieve weight, and sieve estimate
This section deforms the variational data and the sieve estimates of the standard simplex by a parameter \(\varepsilon\) with \(0 \le \varepsilon < 1\). Every object below is defined for all such \(\varepsilon\) and reduces to its standard counterpart at \(\varepsilon = 0\). At \(\varepsilon = 0\) the estimates are the standard ones; at \(\varepsilon > 0\) they give a strictly larger variational problem at the cost of one extra hypothesis on the level of distribution.
The deformation applies two independent scalings to two different pieces of the data. The region on which the profile \(F\) may live is enlarged from \(\{\sum_i t_i \le 1\}\) to \(\{\sum_i t_i \le 1+\varepsilon\}\), which increases the numerator of the Rayleigh quotient by admitting more competitors. Simultaneously, the region over which the \(m\)-th marginal is integrated is shrunk from \(\{\sum_{i \ne m} t_i \le 1\}\) to \(\{\sum_{i \ne m} t_i \le 1-\varepsilon\}\), which decreases it. The reason the two scalings are tied to the same \(\varepsilon\) is arithmetic, not geometric, and is the content of Lemma 1 below: the enlargement pushes the sieve divisors out to \(R^{1+\varepsilon}\), which by itself would carry the moduli of the second-moment bilinear form past the Bombieri–Vinogradov threshold, and the shrinking is exactly the compensating restriction that brings them back.
The enlarged simplex and the shrunken slice
Both regions are corner regions of the same shape as \(\mathcal{R}_k\) (Definition 1), with the total-mass bound \(1\) replaced by \(1+\varepsilon\) and \(1-\varepsilon\) respectively; at \(\varepsilon = 0\) they are \(\mathcal{R}_k\) and \(\mathcal{R}_{k-1}\) exactly. Both are compact, hence Lebesgue measurable and of finite measure, so every integral written below over \(\mathcal{T}_\varepsilon\) or \(\mathcal{S}_\varepsilon\) of a continuous integrand converges absolutely.
Lemma 1 (Enlarged simplex is a homothety of the standard one). For every \(k \ge 2\) and every \(\varepsilon \ge 0\), \(\ \mathcal{T}_\varepsilon = (1+\varepsilon)\cdot\mathcal{R}_k = \{(1+\varepsilon)\,u : u \in \mathcal{R}_k\}\).
Uses: def_simplex
Proof. Write \(c := 1+\varepsilon > 0\). If \(u \in \mathcal{R}_k\) then \(cu\) has non-negative coordinates and \(\sum_i (cu)_i = c\sum_i u_i \le c\cdot 1 = 1+\varepsilon\), so \(cu \in \mathcal{T}_\varepsilon\). Conversely, if \(t \in \mathcal{T}_\varepsilon\) put \(u := c^{-1}t\); then \(u_i = c^{-1}t_i \ge 0\) for each \(i\) and \(\sum_i u_i = c^{-1}\sum_i t_i \le c^{-1}(1+\varepsilon) = 1\), so \(u \in \mathcal{R}_k\) and \(t = cu\). The two inclusions give the claimed equality. ◻
The same argument with \(c := 1-\varepsilon > 0\) gives \(\mathcal{S}_\varepsilon = (1-\varepsilon)\cdot\mathcal{R}_{k-1}\), which is used without further comment below.
Definition 1 (Enlarged profile class). For \(k \ge 2\) and \(0 \le \varepsilon < 1\), the enlarged profile class \(\mathcal{F}_\varepsilon\) is the set of \(F \in C^\infty(\mathbb{R}^k;\mathbb{R})\) with \(\mathop{\mathrm{supp}}F \subseteq \mathcal{T}_\varepsilon\).
Uses: def_simplex
Members of \(\mathcal{F}_\varepsilon\) are automatically compactly supported, since \(\mathcal{T}_\varepsilon\) is compact; and \(\mathcal{F}_0\) is exactly the class of smooth profiles supported on \(\mathcal{R}_k\) used for the standard sieve. No symmetry or sign condition is imposed on \(F\).
Definition 1 (Rescaling operator). For \(\varepsilon \ge 0\) and \(F : \mathbb{R}^k \to \mathbb{R}\), define \(G_\varepsilon F : \mathbb{R}^k \to \mathbb{R}\) by \(\ (G_\varepsilon F)(u) := F\bigl((1+\varepsilon)\,u\bigr)\).
The operator \(G_\varepsilon\) reduces the enlarged development to the standard one. By Lemma 1, \(G_\varepsilon\) maps \(\mathcal{F}_\varepsilon\) bijectively onto \(\mathcal{F}_0\): if \(\mathop{\mathrm{supp}}F \subseteq \mathcal{T}_\varepsilon\) and \((G_\varepsilon F)(u) \ne 0\) then \((1+\varepsilon)u \in \mathcal{T}_\varepsilon = (1+\varepsilon)\cdot\mathcal{R}_k\), whence \(u \in \mathcal{R}_k\); smoothness is preserved because \(G_\varepsilon\) is precomposition with a linear map. On the arithmetic side \(G_\varepsilon\) absorbs the change of the normalising truncation from \(\log R\) to \((1+\varepsilon)\log R\) (Lemma 1); on the analytic side it turns integrals over \(\mathcal{T}_\varepsilon\) into integrals over \(\mathcal{R}_k\) at the price of an explicit Jacobian, as follows.
Lemma 1 (Change of variables on the enlarged simplex). For every \(\varepsilon \ge 0\) and every continuous \(F : \mathbb{R}^k \to \mathbb{R}\), \[\int_{\mathcal{T}_\varepsilon} F(t)^2 \, dt \;=\; (1+\varepsilon)^k \int_{\mathcal{R}_k} \bigl(G_\varepsilon F\bigr)(u)^2 \, du .\]
Uses: def_simplex, def_rescale
Proof. By Lemma 1 the substitution \(t = (1+\varepsilon)u\) maps \(\mathcal{R}_k\) bijectively onto \(\mathcal{T}_\varepsilon\). It is the restriction of the linear map \(u \mapsto (1+\varepsilon)u\) of \(\mathbb{R}^k\), whose determinant is \((1+\varepsilon)^k\); the scaling law for Lebesgue measure on \(\mathbb{R}^k\) therefore gives, for any non-negative measurable \(H\), \(\int_{(1+\varepsilon)\cdot\mathcal{R}_k} H(t)\,dt = (1+\varepsilon)^k \int_{\mathcal{R}_k} H\bigl((1+\varepsilon)u\bigr)\,du\). Apply this with \(H = F^2\) and use \(F\bigl((1+\varepsilon)u\bigr)^2 = (G_\varepsilon F)(u)^2\). ◻
Proof uses: lem_enl_homothety
The enlarged functionals
The denominator of the variational problem is not deformed. For \(F \in \mathcal{F}_\varepsilon\) the quantity \(\int_{\mathcal{T}_\varepsilon}F^2\) equals \(\int_{[0,\infty)^k}F^2 = I_k(F)\), because \(F\) vanishes off \(\mathcal{T}_\varepsilon\); so \(I_k\) of Definition 1 serves unchanged for every \(\varepsilon\), and the deformation enters only through the widened class of competitors. The numerator, by contrast, is deformed: the outer integration is over the shrunken slice.
Lemma 1 (The enlarged marginal in rescaled coordinates). Let \(0 \le \varepsilon < 1\), let \(F \in \mathcal{F}_\varepsilon\), and let \(m \in \{1,\dotsc,k\}\). Then \[J^{(m)}_\varepsilon(F) \;=\; (1+\varepsilon)^{k+1} \int_{\mathcal{R}_{k-1}} \mathbf 1\Bigl[ \textstyle\sum_{i \ne m} u_i \le \frac{1-\varepsilon}{1+\varepsilon} \Bigr] \Bigl( \int_0^\infty (G_\varepsilon F)(u)\, du_m \Bigr)^{\!2} \, du_1 \dotsm du_{m-1}\,du_{m+1}\dotsm du_k .\]
Uses: def_J_k, def_rescale, def_prof_eps, def_simplex
Proof. Substitute \(t = (1+\varepsilon)u\) in both the inner and the outer integral. In the inner integral over the single variable \(t_m\) this contributes one factor \(1+\varepsilon\), so \(\int_0^\infty F(t)\,dt_m = (1+\varepsilon)\int_0^\infty (G_\varepsilon F)(u)\,du_m\); squaring gives \((1+\varepsilon)^2\). In the outer integral over the \(k-1\) variables \((t_i)_{i\ne m}\) it contributes \((1+\varepsilon)^{k-1}\). Together these give the factor \((1+\varepsilon)^{k+1}\).
It remains to identify the domain. Under \(t = (1+\varepsilon)u\) the constraint \((t_i)_{i\ne m} \in \mathcal{S}_\varepsilon\), that is \(t_i \ge 0\) for \(i\ne m\) and \(\sum_{i\ne m} t_i \le 1-\varepsilon\), becomes \(u_i \ge 0\) for \(i \ne m\) and \(\sum_{i\ne m} u_i \le (1-\varepsilon)/(1+\varepsilon)\). Since \((1-\varepsilon)/(1+\varepsilon) \le 1\), this domain is the subset of \(\mathcal{R}_{k-1}\) cut out by the displayed indicator, which is the stated right-hand side. ◻
Lemma 1 exhibits the enlarged marginal as a restricted standard marginal: after rescaling by \(G_\varepsilon\) the profile lives on \(\mathcal{R}_k\), and the only remaining trace of the enlargement is the sharp cutoff \(\sum_{i\ne m}u_i \le (1-\varepsilon)/(1+\varepsilon)\) on the outer variables. That cutoff is the image under the \(\lambda \mapsto y\) transform of the divisor restriction \(\prod_{i \ne m} d_i \le R^{1-\varepsilon}\) imposed in Definition 1.
Lemma 1 (Recovery of the standard data at \(\varepsilon = 0\)). Let \(F\) be compactly supported and Riemann-integrable with \(\mathop{\mathrm{supp}}F \subseteq \mathcal{R}_k\). Then \(J^{(m)}_0(F) = J^{(m)}_k(F)\) for every \(m \in \{1,\dotsc,k\}\).
Uses: def_J_k, def_simplex
Proof. By Definition 1 with \(\varepsilon = 0\) we have \(\mathcal{S}_0 = \mathcal{R}_{k-1} = \mathcal{R}_k^{(m)}\), so the two outer domains agree. The two inner integrands also agree: the standard \(J^{(m)}_k\) integrates \(F\) in the variable \(t_m\) over \([0,\infty)\), and so does \(J^{(m)}_0\). (The hypothesis \(\mathop{\mathrm{supp}}F \subseteq \mathcal{R}_k\) is what makes this identification legitimate: it forces \(F\) to vanish for \(t_m\) outside a bounded range, so both inner integrals converge and are computed by the same one-dimensional integral. Without a support hypothesis the value of \(J^{(m)}_0\) genuinely depends on the values of \(F\) off \(\mathcal{R}_k\), which the standard marginal does not see.) The outer integrals therefore have equal integrands on equal domains. ◻
Proof uses: def_simplex
Consequently, for such \(F\) with \(I_k(F) \ne 0\) one has \(\mathcal{Q}_0(F) = \bigl(\sum_m J^{(m)}_k(F)\bigr)/I_k(F)\): at \(\varepsilon = 0\) the enlarged witness ratio is the standard Rayleigh quotient, and the variational problem introduced next is the standard one.
The supremum is finite: an application of the Cauchy–Schwarz inequality in the variable \(t_m\), whose range is contained in an interval of length \(1+\varepsilon\), gives \(J^{(m)}_\varepsilon(F) \le 2(1+\varepsilon)\|F\|_{L^2(\mathcal{T}_\varepsilon)}^2\) for each \(m\), whence \(0 \le M_{k,\varepsilon} \le 2(1+\varepsilon)k\). Taking \(\varepsilon = 0\), Lemma 1 identifies \(M_{k,0}\) with \(M_k\) of Definition 1, and the bound with \(M_k \le 2k\). In this notation the criterion to be met is \(M_{k,\varepsilon} > 4\) for a single pair \((k, \varepsilon)\); the first case meets it at \((105, 0)\) and the second at \((50, \tfrac1{25})\).
Throughout what follows, \(M_{k,\varepsilon}\) plays a purely nominal role. Every estimate is stated and used in terms of the ratio \(\mathcal{Q}_\varepsilon(F)\) of a concrete competitor \(F\); the supremum is never evaluated, and appears only so that the criterion can be phrased as a property of \((k,\varepsilon)\) alone.
Lemma 1 (A witness certifies the enlarged constant). Let \(0 \le \varepsilon < 1\) and suppose \(F \in L^2(\mathcal{T}_\varepsilon)\) satisfies \[4\,\|F\|_{L^2(\mathcal{T}_\varepsilon)}^2 \;<\; \sum_{m=1}^{k} J^{(m)}_\varepsilon(F) .\] Then \(M_{k,\varepsilon} > 4\).
Proof. Each \(J^{(m)}_\varepsilon\) is non-negative, being an integral of a square; so the right-hand side is non-negative, and the strict inequality forces \(\|F\|_{L^2(\mathcal{T}_\varepsilon)} \ne 0\) (otherwise the left side would be \(0\) and the right side would be \(0\) as well, since \(J^{(m)}_\varepsilon\) vanishes on the zero element of \(L^2\)). Dividing the hypothesis by \(\|F\|^2_{L^2(\mathcal{T}_\varepsilon)} > 0\) gives \(4 < \sum_m J^{(m)}_\varepsilon(F)/\|F\|^2\), and the right-hand side is at most \(M_{k,\varepsilon}\) because \(F\) is one of the competitors in Definition 1. ◻
Lemma 1 (\(L^2\)-continuity of the enlarged witness ratio). Let \(0 \le \varepsilon < 1\), let \(F_1 \in L^2(\mathcal{T}_\varepsilon)\) with \(I_k(F_1) \ne 0\), and let \(\eta > 0\). Then there is \(\zeta > 0\) such that every \(F_2 \in L^2(\mathcal{T}_\varepsilon)\) with \(\|F_2 - F_1\|_{L^2(\mathcal{T}_\varepsilon)} < \zeta\) satisfies \(I_k(F_2) \ne 0\) and \(\mathcal{Q}_\varepsilon(F_2) > \mathcal{Q}_\varepsilon(F_1) - \eta\).
Uses: def_M_k, def_simplex
Proof. On \(L^2(\mathcal{T}_\varepsilon)\) the denominator of \(\mathcal{Q}_\varepsilon\) is \(F \mapsto \|F\|_{L^2(\mathcal{T}_\varepsilon)}^2\), which is continuous (Lemma 1), and the numerator is \(F \mapsto \sum_m J^{(m)}_\varepsilon(F)\), a finite sum of continuous functionals (Lemma 1 applied on \(\mathcal{T}_\varepsilon\); each \(J^{(m)}_\varepsilon\) is the squared norm of the composite of the partial-integration map \(L^2(\mathcal{T}_\varepsilon) \to L^2(\mathbb{R}^{k-1})\) with the restriction map onto \(\mathcal{S}_\varepsilon\), both bounded linear). A quotient of continuous real functions is continuous at every point where the denominator is non-zero; since \(\|F_1\|^2 = I_k(F_1) \ne 0\), the map \(\mathcal{Q}_\varepsilon\) is continuous at \(F_1\). Continuity at \(F_1\) furnishes \(\zeta_1 > 0\) with \(|\mathcal{Q}_\varepsilon(F_2) - \mathcal{Q}_\varepsilon(F_1)| < \eta\) whenever \(\|F_2 - F_1\| < \zeta_1\); shrinking \(\zeta_1\) further so that \(\|F_2 - F_1\| < \zeta\) forces \(\|F_2\| \ge \|F_1\| - \zeta > 0\), hence \(I_k(F_2) \ne 0\), yields the claim. ◻
Proof uses: lem_Ik_L2_continuous, lem_Jk_m_L2_continuous
The certificate-admissible level
Two of the estimates below — the aggregate Chinese remainder error of Lemma 1 and the Bombieri–Vinogradov modulus bound of Lemma 1 — require the enlarged divisor range \(R^{1+\varepsilon}\) to stay within the ranges those estimates tolerate. Both requirements, together with the requirement that the resulting sieve criterion be satisfiable at all, are collected in the following single condition on the level \(\vartheta\).
Definition 1 (Certificate-admissible level). Let \(\varepsilon \ge 0\) and \(Q \in \mathbb{R}\). A real number \(\vartheta\) is \((\varepsilon, Q)\)-certificate-admissible if
\(0 < \vartheta < \tfrac12\);
\(2/\vartheta < Q\);
\(1 + \varepsilon < 1/\vartheta\);
and the primes have level of distribution \(\vartheta\).
Both \(\varepsilon\) and the threshold \(Q\) are parameters of the definition, not constants: nothing in the theory below depends on a particular numerical value of either.
The three clauses are consumed in three different places.
Clause (iii), equivalently \((1+\varepsilon)\vartheta < 1\), is the enlargement room. It makes the enlarged truncation harmless in the first moment: with \(R = N^{\vartheta/2 - \delta}\) one has \(R^{2(1+\varepsilon)} = N^{(1+\varepsilon)(\vartheta - 2\delta)}\), and (iii) gives \((1+\varepsilon)(\vartheta - 2\delta) < (1+\varepsilon)\vartheta < 1\), so the aggregate error of Lemma 1 is \(o(N)\) and Proposition 1 needs no equidistribution input at all. Clause (iii) also guarantees \((1+\varepsilon)(\vartheta/2 - \delta) < \tfrac12\), i.e. \(R^{1+\varepsilon} < N^{1/2}\), which lets the enlarged weight be rewritten as a standard weight at a legitimate truncation level (Lemma 1).
Clause (i) is the Bombieri–Vinogradov room. The moduli occurring in the second-moment bilinear form are bounded by \(R^{a+b}\) where \(a, b\) are the exponents of the outer divisor products of the two weights; the support split of Definition 1 arranges \(a + b \le 2\), so the moduli are at most \(R^2 = N^{\vartheta - 2\delta}\), and (i) puts this below \(N^{1/2}\) (Lemma 1). Note that without the split the bound would only be \(R^{2(1+\varepsilon)}\), which (iii) keeps below \(N\) but not below \(N^{1/2}\); the split is therefore the mechanism that makes the scaling compatible with Bombieri–Vinogradov.
Clause (ii) is the sieve criterion. It makes the positivity window of Lemma 1 non-empty, and hence gives the value a numerical \(\mathcal{Q}_\varepsilon\) must exceed. Since \(\vartheta\) may be taken as close to \(\tfrac12\) as one likes — but not equal to it — clause (ii) is satisfiable for some \(\vartheta\) exactly when \(Q > 4\), which is why the target of the enlarged variational problem is \(M_{k,\varepsilon} > 4\) and not \(M_{k,\varepsilon} \ge 4\).
Clause (iii) is implied by clause (i) whenever \(\varepsilon \le 1\): then \(1+\varepsilon \le 2 < 1/\vartheta\). It is nevertheless stated separately because it is the clause the scaling consumes, and because it is the one that would bind first if \(\varepsilon\) were pushed towards or beyond \(1\).
The enlarged sieve weight and its support
Definition 1 (\(\varepsilon\)-permissible support). Let \(\varepsilon \ge 0\). A function \(\lambda : \mathbb{N}^k \to \mathbb{R}\) has \(\varepsilon\)-permissible support if \(\lambda(d_1,\dotsc,d_k) = 0\) whenever one of the following fails:
\(d_i \ge 1\) for every \(i\), and \(\prod_{i=1}^{k} d_i \le R^{1+\varepsilon}\);
\(\gcd\bigl(\prod_{i=1}^{k} d_i,\, W\bigr) = 1\);
\(\prod_{i=1}^{k} d_i\) is squarefree.
Uses: def_R, def_W_trick, def_adm_support
This is Definition 1 with the single bound \(R\) relaxed to \(R^{1+\varepsilon}\); at \(\varepsilon = 0\) the two conditions coincide.
Definition 1 (Enlarged sieve weight). Let \(\varepsilon \ge 0\) and \(F \in \mathcal{F}_\varepsilon\). Define \(\lambda_\varepsilon : \mathbb{N}^k \to \mathbb{R}\) by \[\lambda_\varepsilon(d_1,\dotsc,d_k) \;:=\; \Bigl( \prod_{i=1}^{k} \mu(d_i)\, d_i \Bigr) \sum_{\substack{r_1, \dotsc, r_k \\ d_i \mid r_i \ \forall i \\ \prod_i r_i \le R^{1+\varepsilon} \\ \gcd(\prod_i r_i, W) = 1,\ \mu(\prod_i r_i)^2 = 1}} \frac{1}{\prod_{i=1}^{k} \varphi(r_i)} \, F\Bigl( \frac{\log r_1}{\log R}, \dotsc, \frac{\log r_k}{\log R} \Bigr) .\]
Uses: def_prof_eps, def_R, def_W_trick, def_lambda_from_F
Two features of this definition are used below, and neither is forced by the shape of the formula. First, the \(\mathbf r\)-sum ranges over the enlarged region \(\prod_i r_i \le R^{1+\varepsilon}\) rather than \(\prod_i r_i \le R\): this is the only change relative to the standard \(F\)-derived weight of Definition 1. Second, the argument of \(F\) is normalised by \(\log R\) with the original \(R = N^{\vartheta/2-\delta}\), not by \(\log R^{1+\varepsilon}\). Consequently the point \(\bigl(\tfrac{\log r_1}{\log R}, \dotsc, \tfrac{\log r_k}{\log R}\bigr)\) lies in \(\mathcal{T}_\varepsilon\) exactly when \(\prod_i r_i \le R^{1+\varepsilon}\): the sieve range and the support of \(F\) match, so that \(\mathcal{T}_\varepsilon\) is the correct region.
Lemma 1 (The enlarged weight is \(\varepsilon\)-permissible). For every \(\varepsilon \ge 0\) and every \(F \in \mathcal{F}_\varepsilon\), the weight \(\lambda_\varepsilon\) has \(\varepsilon\)-permissible support.
Proof. Fix \(\mathbf d\) with \(\lambda_\varepsilon(\mathbf d) \ne 0\). If some \(d_i = 0\) then the factor \(\prod_i \mu(d_i)d_i\) vanishes, contradiction; so \(d_i \ge 1\) for all \(i\). Since \(\lambda_\varepsilon(\mathbf d) \ne 0\), the defining sum is non-empty, so there is at least one \(\mathbf r\) with \(d_i \mid r_i\) for every \(i\) satisfying the three summation constraints. From \(\prod_i r_i \le R^{1+\varepsilon}\) and \(d_i \mid r_i\) (hence \(d_i \le r_i\)) we get \(\prod_i d_i \le \prod_i r_i \le R^{1+\varepsilon}\), which is clause (i). From \(\gcd(\prod_i r_i, W) = 1\) and \(\prod_i d_i \mid \prod_i r_i\) we get \(\gcd(\prod_i d_i, W) = 1\), which is clause (ii). From \(\prod_i r_i\) squarefree and \(\prod_i d_i \mid \prod_i r_i\) we get \(\prod_i d_i\) squarefree, which is clause (iii). ◻
Lemma 1 (The enlarged weight is a standard weight at a shifted level). Let \(\varepsilon \ge 0\), let \(F \in \mathcal{F}_\varepsilon\), and let \(\vartheta, \delta\) satisfy \(0 < \vartheta < 1\), \(0 < \delta < \vartheta/2\) and \(1+\varepsilon < 1/\vartheta\). Let \(\Theta\) and \(\Delta\) be the real numbers satisfying \[\Theta = \tfrac12 + (1+\varepsilon)\bigl(\tfrac{\vartheta}{2} - \delta\bigr), \qquad \Delta = \tfrac14 - \tfrac12 (1+\varepsilon)\bigl(\tfrac{\vartheta}{2} - \delta\bigr) .\] Then \(0 < \Theta < 1\), \(0 < \Delta < \Theta/2\), the truncation attached to \((\Theta, \Delta)\) is \(N^{\Theta/2 - \Delta} = R^{1+\varepsilon}\), and \(\lambda_\varepsilon\) is the standard \(F\)-derived weight of Definition 1 at truncation \(R^{1+\varepsilon}\) applied to the rescaled profile \(G_\varepsilon F\).
Proof. Write \(\alpha := (1+\varepsilon)(\vartheta/2 - \delta)\), so \(\Theta = \tfrac12 + \alpha\) and \(\Delta = \tfrac14 - \tfrac{\alpha}{2}\), and hence \(\Theta/2 - \Delta = \alpha\) identically. We have \(\alpha > 0\) because \(\delta < \vartheta/2\); and \(\alpha < (1+\varepsilon)\vartheta/2 < \tfrac12\) by \(1+\varepsilon < 1/\vartheta\). From \(0 < \alpha < \tfrac12\) the two parameter constraints \(0 < \Theta < 1\) and \(0 < \Delta < \Theta/2\) follow by inspection. Also \(N^{\Theta/2 - \Delta} = N^{\alpha} = \bigl(N^{\vartheta/2 - \delta}\bigr)^{1+\varepsilon} = R^{1+\varepsilon}\).
For the last assertion, compare Definition 1 with Definition 1 at truncation \(R' := R^{1+\varepsilon}\). The two Möbius prefactors and the two summation ranges agree verbatim, the latter because the standard weight at truncation \(R'\) sums over \(\prod_i r_i \le R' = R^{1+\varepsilon}\) subject to the same coprimality and squarefreeness conditions. Only the argument of the profile differs: Definition 1 at truncation \(R'\) evaluates the profile at \(\bigl(\tfrac{\log r_i}{\log R'}\bigr)_i\), whereas Definition 1 evaluates \(F\) at \(\bigl(\tfrac{\log r_i}{\log R}\bigr)_i\). Since \(\log R' = (1+\varepsilon)\log R\), we have \(\bigl(\tfrac{\log r_i}{\log R}\bigr)_i = (1+\varepsilon)\bigl(\tfrac{\log r_i}{\log R'}\bigr)_i\), so by Definition 1 \[F\Bigl( \tfrac{\log r_1}{\log R}, \dotsc, \tfrac{\log r_k}{\log R} \Bigr) \;=\; (G_\varepsilon F)\Bigl( \tfrac{\log r_1}{\log R'}, \dotsc, \tfrac{\log r_k}{\log R'} \Bigr),\] which is exactly the profile value the standard weight at truncation \(R'\) would use for \(G_\varepsilon F\). ◻
Lemma 1 shows the enlarged sieve weight is not a new object but the standard weight, evaluated at a different truncation level and applied to a rescaled profile. Every estimate of the standard development that is uniform in the level parameters therefore transfers to \(\lambda_\varepsilon\) without new work, and the only new arguments are those where the enlargement meets a fixed threshold — the Bombieri–Vinogradov threshold in the second moment.
The enlarged \(S_1\) asymptotic
Lemma 1 (Aggregate CRT error for \(\varepsilon\)-permissible weights). Let \(\varepsilon \ge 0\) and let \(\lambda, \lambda'\) have \(\varepsilon\)-permissible support with \(|\lambda|, |\lambda'| \le 1\) pointwise, and suppose \((1+\varepsilon)\vartheta < 1\). Then \[\begin{gathered} \Biggl| \sum_{\mathbf d, \mathbf e} \lambda(\mathbf d)\, \lambda'(\mathbf e) \Biggl( \sum_{\substack{N < n \le 2N \\ n \equiv v_0 \ (\mathrm{mod}\ W) \\ [d_i, e_i] \mid n + h_i\ \forall i}} 1 \;-\; \frac{N}{W \prod_i [d_i, e_i]}\, \mathbf 1\bigl[W, [d_1,e_1], \dotsc, [d_k,e_k] \text{ pairwise coprime}\bigr] \Biggr) \Biggr| \\ \ll_k\ R^{2(1+\varepsilon)} (\log R)^{2(k-1)} \;=\; o(N) . \end{gathered}\]
Uses: def_eps_adm_support, def_R
Proof. Write \(\rho(\mathbf d, \mathbf e)\) for the parenthesised difference. By Lemma 1 the inner counting sum equals \(N/(W\prod_i[d_i,e_i]) + O(1)\) when \(W, [d_1,e_1], \dotsc, [d_k,e_k]\) are pairwise coprime and equals \(0\) otherwise, so in either branch \(|\rho(\mathbf d, \mathbf e)| \le C_k\) for a constant depending only on \(k\). With \(|\lambda|, |\lambda'| \le 1\) this gives \[\Bigl| \sum_{\mathbf d, \mathbf e} \lambda(\mathbf d) \lambda'(\mathbf e) \rho(\mathbf d, \mathbf e) \Bigr| \;\le\; C_k \cdot \#\bigl\{ (\mathbf d, \mathbf e) : \lambda(\mathbf d)\lambda'(\mathbf e) \ne 0 \bigr\} .\] By \(\varepsilon\)-permissibility each factor is supported on tuples with \(\prod_i d_i \le R^{1+\varepsilon}\) and \(\prod_i d_i\) squarefree. Writing \(D = \prod_i d_i\), the number of ordered \(k\)-tuples with \(\prod_i d_i = D\) is \(\tau_k(D)\), so the number of admissible \(\mathbf d\) is at most \(\sum_{D \le R^{1+\varepsilon}} \tau_k(D) \ll_k R^{1+\varepsilon}(\log R)^{k-1}\) by the standard divisor-sum bound; squaring gives the pair count.
Finally \(R = N^{\vartheta/2 - \delta}\) gives \(R^{2(1+\varepsilon)} = N^{(1+\varepsilon)(\vartheta - 2\delta)}\), and \((1+\varepsilon)(\vartheta - 2\delta) < (1+\varepsilon)\vartheta < 1\); the logarithmic factor is \(N^{o(1)}\), so the whole bound is \(o(N)\). ◻
Proof uses: lem_S1_CRT
Proposition 1 (Enlarged \(S_1\) asymptotic). Let \(k \ge 2\), let \(\mathcal{H}\) be admissible, let \(\varepsilon \ge 0\), and let \(F \in \mathcal{F}_\varepsilon\). Let \(0 < \vartheta < 1\), \(0 < \delta < \vartheta/2\) and \(1 + \varepsilon < 1/\vartheta\). Then, with \(R = N^{\vartheta/2-\delta}\) the unenlarged truncation, \[S_1(\lambda_\varepsilon) \;=\; \bigl(1 + O(D_0^{-1})\bigr)\, \frac{\varphi(W)^k\, N\, (\log R)^k}{W^{k+1}} \int_{\mathcal{T}_\varepsilon} F(t)^2 \, dt\] as \(N \to \infty\), the implied constant depending only on \(k\), \(\mathcal{H}\) and \(F_{\max}(G_\varepsilon F)\).
Uses: def_lambda_eps, def_S1, def_prof_eps, def_simplex, def_R
Proof. By Lemma 1 the parameters \((\Theta, \Delta)\) are legitimate level parameters and \(\lambda_\varepsilon\) is the standard \(F\)-derived weight at truncation \(R' = R^{1+\varepsilon} = N^{\Theta/2-\Delta}\) applied to \(G_\varepsilon F\). The rescaled profile \(G_\varepsilon F\) is smooth and, by Lemma 1, supported on \(\mathcal{R}_k\); it is therefore an admissible input to the standard smooth first-moment asymptotic Lemma 1, whose hypotheses involve only \(k\), \(\mathcal{H}\) and the level parameters. That lemma gives \[S_1(\lambda_\varepsilon) \;=\; \bigl(1 + O(D_0^{-1})\bigr)\, \frac{\varphi(W)^k\, N\, (\log R')^k}{W^{k+1}} \int_{\mathcal{R}_k} \bigl(G_\varepsilon F\bigr)(u)^2\, du .\] (The equidistribution input needed here is only the aggregate CRT error of Lemma 1, which is \(o(N)\) by \((1+\varepsilon)\vartheta < 1\); no Bombieri–Vinogradov estimate is used in the first moment.)
It remains to restore the original normalisation. Since \(\log R' = (1+\varepsilon)\log R\) we have \((\log R')^k = (1+\varepsilon)^k (\log R)^k\), and by Lemma 1 \(\int_{\mathcal{R}_k}(G_\varepsilon F)^2 = (1+\varepsilon)^{-k}\int_{\mathcal{T}_\varepsilon}F^2\). The two factors \((1+\varepsilon)^{\pm k}\) cancel, and the same cancellation applies to the error term, giving the stated form with the unenlarged \(\log R\) and the integral over \(\mathcal{T}_\varepsilon\). ◻
The cancellation of \((1+\varepsilon)^{\pm k}\) makes the enlarged first moment identical in form to the standard one, with \(I_k(F)\) now computed over the larger region. Since \(I_k\) is the denominator of the variational problem, enlarging the simplex costs nothing in the denominator; the entire effect of the enlargement is in the numerator, treated next.
The support split and the enlarged \(S_2\) lower bound
Definition 1 (Support-split weight). Fix \(m \in \{1,\dotsc,k\}\) and \(\varepsilon \ge 0\). Set \[\lambda_{1,m}(\mathbf d) \;:=\; \lambda_\varepsilon(\mathbf d)\, \mathbf 1\Bigl[ \textstyle\prod_{i \ne m} d_i \le R^{1-\varepsilon} \Bigr] .\]
Uses: def_lambda_eps, def_R
Definition 1 (Complementary weight). With the notation of Definition 1, set \(\lambda_{2,m} := \lambda_\varepsilon - \lambda_{1,m}\).
Uses: def_lambda_eps, def_split
Explicitly, \(\lambda_{2,m}(\mathbf d) = \lambda_\varepsilon(\mathbf d)\) when \(\prod_{i\ne m} d_i > R^{1-\varepsilon}\) and \(0\) otherwise; in particular \(\lambda_{2,m}\) inherits from \(\lambda_\varepsilon\) the outer bound \(\prod_{i\ne m} d_i \le \prod_i d_i \le R^{1+\varepsilon}\). The two outer exponents are therefore \(1-\varepsilon\) for \(\lambda_{1,m}\) and \(1+\varepsilon\) for \(\lambda_{2,m}\), and the point of the split is the identity \((1-\varepsilon) + (1+\varepsilon) = 2\).
Notation 1 (Split divisor sums). For \(\ell \in \{1,2\}\) and \(n\) in the sieve range, set \[A_{\ell,m}(n) \;:=\; \sum_{\substack{\mathbf d \\ d_i \mid n + h_i\ \forall i}} \lambda_{\ell,m}(\mathbf d),\] so that the inner divisor sum defining \(w_n\) is \(A_{1,m}(n) + A_{2,m}(n)\).
Uses: def_split, def_split2, def_weight
Lemma 1 (Pointwise lower bound). For every \(n\) in the sieve range, \(\ w_n \ge A_{1,m}(n)^2 + 2\,A_{1,m}(n)\,A_{2,m}(n)\).
Uses: not_a, def_weight
Proof. Put \(x_1 := A_{1,m}(n)\) and \(x_2 := A_{2,m}(n)\). By Notation 1, \(w_n = (x_1 + x_2)^2 = x_1^2 + 2x_1x_2 + x_2^2 \ge x_1^2 + 2x_1x_2\), since \(x_2^2 \ge 0\). ◻
Lemma 1 discards the one term, \(A_{2,m}(n)^2\), that the scaling makes unavailable: it is a bilinear form in \((\lambda_{2,m}, \lambda_{2,m})\), whose outer exponents sum to \(2(1+\varepsilon) > 2\), and whose moduli therefore exceed the Bombieri–Vinogradov threshold. The two terms that are kept have outer exponent sums \(2(1-\varepsilon) \le 2\) and \((1-\varepsilon) + (1+\varepsilon) = 2\) respectively, and are handled by the following bound.
Lemma 1 (Bombieri–Vinogradov modulus bound). Fix \(m \in \{1,\dotsc,k\}\) and \(A > 0\). Let \(\vartheta\) satisfy Definition 1(i) and (iii) and have level of distribution \(\vartheta\), and let \(a, b \ge 0\) with \(a + b \le 2\). Let \(\lambda, \lambda'\) have \(\varepsilon\)-permissible support and be supported additionally on \(\{\prod_{i \ne m} d_i \le R^{a}\}\) and \(\{\prod_{i \ne m} d_i \le R^{b}\}\) respectively. Then \[\sum_{N < n \le 2N} \chi_{\mathbb{P}}(n + h_m) \sum_{\mathbf d, \mathbf e} \lambda(\mathbf d)\, \lambda'(\mathbf e) \prod_{i=1}^{k} \mathbf 1\bigl[ [d_i, e_i] \mid n + h_i \bigr]\] differs from its expected main term by \(O_A\bigl(N/(\log N)^A\bigr)\).
Uses: def_eps_adm_support, def_cert_adm, def_R
Proof. The factor \(\chi_{\mathbb{P}}(n+h_m)\) restricts \(n\) to integers with \(n + h_m\) prime, so any divisor \(d_m\) of \(n + h_m\) lies in \(\{1,\, n+h_m\}\). By \(\varepsilon\)-permissibility, \(d_m \le \prod_i d_i \le R^{1+\varepsilon} = N^{(1+\varepsilon)(\vartheta/2-\delta)}\), and clause (iii) of Definition 1 gives \((1+\varepsilon)(\vartheta/2-\delta) < (1+\varepsilon)\vartheta/2 < \tfrac12\), so \(d_m < N^{1/2} < N < n + h_m\) for large \(N\); hence \(d_m = 1\), and likewise \(e_m = 1\). The \(m\)-th coordinate therefore contributes only the trivial divisor, and the effective modulus of the bilinear form is \(q = \prod_{i \ne m} [d_i, e_i]\).
For \((\mathbf d, \mathbf e)\) in the joint support, \([d_i, e_i] \le d_i e_i\) and the extra outer support hypotheses give \(q \le \prod_{i\ne m} d_i \cdot \prod_{i \ne m} e_i \le R^{a} R^{b} = R^{a+b} \le R^{2} = N^{\vartheta - 2\delta}\); this is Lemma 1 with \(B_1 = R^{a}\) and \(B_2 = R^{b}\), whose point is precisely that the modulus depends on the two outer truncations only through their product, so that \(a + b \le 2\) suffices even when one of \(a\), \(b\) exceeds \(1\). By clause (i) of Definition 1, \(\vartheta < \tfrac12\), so \(\vartheta - 2\delta < \tfrac12\), and Lemma 1 supplies \(\eta_R > 0\) with \(qW \le N^{1/2 - \eta_R}\) for all large \(N\). Every modulus occurring in the bilinear form thus lies below the Bombieri–Vinogradov threshold, and the Cauchy–Schwarz-plus-Bombieri–Vinogradov bilinear estimate of Lemma 1 bounds the aggregate error by \(O_A(N/(\log N)^A)\) for every \(A > 0\). (The pointwise divisor-weight cost of size \((\log N)^{O(k)}\) incurred in applying that estimate is absorbed by requesting the exponent \(A + 2k\) in place of \(A\) from the same level-of-distribution hypothesis.) ◻
Proof uses: lem_R2W_below_N_half_minus_eps, lem_S2m_split_modulus, lem_S2m_error
Lemma 1 (Retained bilinear forms have negligible error). Fix \(m \in \{1,\dotsc,k\}\) and \(A > 0\), let \(\vartheta\) satisfy Definition 1(i) and (iii) and have level of distribution \(\vartheta\), and let \(F \in \mathcal{F}_\varepsilon\). Then for each \(\lambda' \in \{\lambda_{1,m},\, \lambda_{2,m}\}\) the sum \[\sum_{N < n \le 2N} \chi_{\mathbb{P}}(n + h_m)\, A_{1,m}(n)\, A'_{m}(n), \qquad A'_m(n) = \sum_{d_i \mid n+h_i \forall i} \lambda'(\mathbf d),\] differs from its expected main term by \(O_A\bigl(N/(\log N)^A\bigr)\).
Uses: def_split, def_split2, not_a, lem_bv_modulus, def_cert_adm
Proof. Expanding the product of the two divisor sums exhibits the displayed quantity as the bilinear form of Lemma 1 with the pair \((\lambda_{1,m}, \lambda')\). Both weights are \(\varepsilon\)-permissible: \(\lambda_{1,m}\) and \(\lambda_{2,m}\) are obtained from \(\lambda_\varepsilon\) by multiplying by an indicator, respectively by subtracting such a product, and \(\varepsilon\)-permissibility (Definition 1) is preserved by both operations; \(\lambda_\varepsilon\) itself is \(\varepsilon\)-permissible by Lemma 1.
It remains to check the outer exponents. By Definition 1, \(\lambda_{1,m}\) is supported on \(\prod_{i\ne m} d_i \le R^{1-\varepsilon}\), so \(a = 1-\varepsilon\). For \(\lambda' = \lambda_{1,m}\) we may take \(b = 1-\varepsilon\), and \(a + b = 2(1-\varepsilon) \le 2\). For \(\lambda' = \lambda_{2,m}\) we may take \(b = 1+\varepsilon\): indeed \(\lambda_{2,m}(\mathbf d) \ne 0\) implies \(\lambda_\varepsilon(\mathbf d) \ne 0\), whence \(\prod_{i \ne m} d_i \le \prod_i d_i \le R^{1+\varepsilon}\); and then \(a + b = (1-\varepsilon) + (1+\varepsilon) = 2\). In both cases \(a + b \le 2\) and Lemma 1 applies. ◻
Proof uses: lem_bv_modulus, lem_lambda_perm, def_eps_adm_support
Lemma 1 (Main term of the retained square). Fix \(m \in \{1,\dotsc,k\}\), let \(\varepsilon \ge 0\), let \(F \in \mathcal{F}_\varepsilon\), and let \(\vartheta\) satisfy Definition 1(i) and (iii) and have level of distribution \(\vartheta\). Then, for every \(A > 0\), \[\sum_{N < n \le 2N} \chi_{\mathbb{P}}(n + h_m)\, A_{1,m}(n)^2 \;=\; \bigl(1 + o(1)\bigr)\, \frac{\varphi(W)^k\, N\, (\log R)^{k+1}}{W^{k+1} \log N}\, J^{(m)}_\varepsilon(F) \;+\; O_A\Bigl( \frac{N}{(\log N)^A} \Bigr) .\]
Uses: def_split, not_a, def_J_k, lem_retained, def_cert_adm
Proof. By Lemma 1 with \(\lambda' = \lambda_{1,m}\), replacing the sum by its arithmetic main term commits an error \(O_A(N/(\log N)^A)\); here the level of distribution enters, through Lemma 1. The identification of that arithmetic main term, for a bilinear expression in two weights rather than a square, is the two-weight Chinese remainder bound of Lemma 1, applied to the pair \((\lambda_{1,m}, \lambda_{1,m})\).
The arithmetic main term is evaluated by the standard \(S_2^{(m)}\) chain, applied to \(\lambda_{1,m}\) regarded — via Lemma 1 — as a standard weight at truncation \(R' = R^{1+\varepsilon}\) built from \(G_\varepsilon F\), and carrying the additional outer restriction \(\prod_{i \ne m} r_i \le R^{1-\varepsilon}\). The chain consists of the extraction of the diagonal form (Lemma 1), the reduction to a second moment (Lemma 1), the expansion and decoupling step (Lemma 1), and the evaluation of the resulting decoupled \((k+1)\)-fold sum (Lemma 1). Each of these uses the divisor range only through the coordinate-wise bound \(\prod_i r_i \le R'\) and its logarithm, and so applies verbatim at the enlarged truncation.
The sharp cutoff \(\prod_{i\ne m} r_i \le R^{1-\varepsilon}\) is not directly admissible in the chain, which requires a smooth profile. It is therefore replaced by a smooth ramp of width \(\eta > 0\) in the logarithmic variables, i.e. \(\lambda_{1,m}\) is replaced by the weight built from the profile \(u \mapsto \chi_\eta\bigl(\sum_{i\ne m} u_i\bigr)\,(G_\varepsilon F)(u)\) where \(\chi_\eta\) is a smooth non-increasing transition equal to \(1\) below \(\frac{1-\varepsilon}{1+\varepsilon} - \eta\) and to \(0\) above \(\frac{1-\varepsilon}{1+\varepsilon}\). This is a legitimate smooth profile supported on \(\mathcal{R}_k\), so the chain applies to it for each fixed \(\eta\), and produces the prefactor \(\varphi(W)^k N (\log R')^{k+1} / (W^{k+1}\log N)\) times the standard marginal \(J^{(m)}_k\) of the ramped rescaled profile. Letting \(\eta \to 0\), dominated convergence in the outer \((k-1)\)-fold integral replaces the ramp by the sharp indicator, so the marginal converges to \[\int_{\mathcal{R}_{k-1}} \mathbf 1\Bigl[ \textstyle\sum_{i\ne m} u_i \le \frac{1-\varepsilon}{1+\varepsilon} \Bigr] \Bigl( \int_0^\infty (G_\varepsilon F)(u)\, du_m \Bigr)^{\!2} du_{\hat m} \;=\; \frac{J^{(m)}_\varepsilon(F)}{(1+\varepsilon)^{k+1}} ,\] the last equality being Lemma 1. Since \((\log R')^{k+1} = (1+\varepsilon)^{k+1}(\log R)^{k+1}\), the factors \((1+\varepsilon)^{\pm(k+1)}\) cancel and the main term takes the displayed form. ◻
Lemma 1 (Smallness of the retained–complementary cross term). In the situation of Lemma 1, for every \(\kappa > 0\) there is a choice of the smoothing width for which \[\Bigl| \sum_{N < n \le 2N} \chi_{\mathbb{P}}(n + h_m)\, A_{1,m}(n)\, A_{2,m}(n) \Bigr| \;\le\; \kappa \cdot \frac{\varphi(W)^k\, N\, (\log R)^{k+1}}{W^{k+1} \log N} \;+\; O_A\Bigl( \frac{N}{(\log N)^A} \Bigr)\] for every \(A > 0\) and all sufficiently large \(N\).
Uses: def_split, def_split2, not_a, lem_retained, lem_s2_a1_main
Proof. By Lemma 1 with \(\lambda' = \lambda_{2,m}\), replacing the sum by its arithmetic main term commits an error \(O_A(N/(\log N)^A)\), and the main term is produced by the same \(S_2^{(m)}\) chain as in Lemma 1, now applied to the pair \((\lambda_{1,m}, \lambda_{2,m})\); in particular it carries the same prefactor \(\varphi(W)^kN(\log R)^{k+1}/(W^{k+1}\log N)\).
What remains is to bound the resulting integral. In the rescaled logarithmic coordinates \(u_i = \log r_i / \log R'\), the profile attached to \(\lambda_{1,m}\) is supported where \(\sum_{i\ne m} u_i \le \frac{1-\varepsilon}{1+\varepsilon}\) and the profile attached to \(\lambda_{2,m} = \lambda_\varepsilon - \lambda_{1,m}\) is supported where \(\sum_{i\ne m} u_i > \frac{1-\varepsilon}{1+\varepsilon}\). The two supports are disjoint — this is the same disjointness that makes the diagonal form additive along the split in Lemma 1, where it leaves no cross term at all — so with the sharp cutoff the product would vanish identically; with the smooth ramp of width \(\eta\) used in Lemma 1, the product is supported on the boundary strip \[\mathcal{B}_\eta \;:=\; \Bigl\{ u \in \mathcal{R}_k : \tfrac{1-\varepsilon}{1+\varepsilon} - \eta \le \textstyle\sum_{i \ne m} u_i \le \tfrac{1-\varepsilon}{1+\varepsilon} \Bigr\} ,\] whose \((k-1)\)-dimensional outer slice has Lebesgue measure at most \(\eta/(k-2)!\). Since \(G_\varepsilon F\) is bounded on the compact set \(\mathcal{R}_k\) and the ramp takes values in \([0,1]\), the integrand is \(O(1)\) on \(\mathcal{B}_\eta\) and the integral is \(O(\eta)\), with an implied constant depending only on \(k\) and \(\sup|G_\varepsilon F|\). Choosing \(\eta\) so small that this is at most \(\kappa\) gives the claim. ◻
Proposition 1 (Enlarged \(S_2^{(m)}\) lower bound). Let \(k \ge 2\), let \(\mathcal{H}\) be admissible, let \(\varepsilon \ge 0\), let \(F \in \mathcal{F}_\varepsilon\), and let \(\vartheta\) satisfy Definition 1(i) and (iii) and have level of distribution \(\vartheta\). Then for each \(m \in \{1,\dotsc,k\}\), \[S_2^{(m)}(\lambda_\varepsilon) \;\ge\; \bigl(1 + o(1)\bigr)\, \frac{\varphi(W)^k\, N\, (\log R)^{k+1}}{W^{k+1} \log N}\, J^{(m)}_\varepsilon(F)\] as \(N \to \infty\).
Uses: def_lambda_eps, def_S2m, def_J_k, def_cert_adm, def_prof_eps
Proof. Multiplying Lemma 1 by \(\chi_{\mathbb{P}}(n + h_m) \ge 0\) and summing over \(N < n \le 2N\) with \(n \equiv v_0 \pmod W\) gives \[S_2^{(m)}(\lambda_\varepsilon) \;\ge\; \sum_n \chi_{\mathbb{P}}(n+h_m) A_{1,m}(n)^2 \;+\; 2 \sum_n \chi_{\mathbb{P}}(n+h_m) A_{1,m}(n) A_{2,m}(n) .\] Write \(P := \varphi(W)^k N (\log R)^{k+1} W^{-k-1}(\log N)^{-1}\) for the common prefactor. By Lemma 1 the first sum is \((1+o(1))\,P\,J^{(m)}_\varepsilon(F) + O_A(N/(\log N)^A)\). By Lemma 1, for any \(\kappa > 0\) the second sum is at most \(2\kappa P + O_A(N/(\log N)^A)\) in absolute value. Given \(\eta > 0\), if \(J^{(m)}_\varepsilon(F) > 0\) choose \(\kappa < \eta J^{(m)}_\varepsilon(F)/2\); then \[S_2^{(m)}(\lambda_\varepsilon) \;\ge\; \bigl(1 - \eta + o(1)\bigr) P\, J^{(m)}_\varepsilon(F) \;+\; O_A\bigl(N/(\log N)^A\bigr) .\] As \(\eta > 0\) was arbitrary the claim follows. If \(J^{(m)}_\varepsilon(F) = 0\) the right-hand side of the claim is \(0\) while the left-hand side is a sum of squares against a non-negative weight, hence non-negative, and the claim is trivial. ◻
Proof uses: lem_pointwise, lem_s2_a1_main, lem_s2_cross, lem_S2_decomp
The criterion in enlarged form
Lemma 1 (The positivity window is non-empty). Let \(0 < \vartheta < \tfrac12\) and \(Q \in \mathbb{R}\) with \(2/\vartheta < Q\). Then there exist \(\delta \in (0, \vartheta/2)\) and \(\rho \in (1,2)\) with \(\bigl(\tfrac{\vartheta}{2} - \delta\bigr) Q > \rho\).
Proof. From \(2/\vartheta < Q\) and \(\vartheta > 0\) we get \(2 < Q\vartheta\), i.e. \(M := \tfrac{\vartheta}{2}Q > 1\). Put \(u := \min(2, M) \in (1, 2]\) and \(\rho := \tfrac{1+u}{2}\); then \(1 < \rho < u \le 2\) and \(\rho < M\), so \(\rho \in (1,2)\) and \(\rho/Q < \vartheta/2\) (note \(Q > 2/\vartheta > 0\)). Finally set \(\delta := \tfrac12\bigl(\tfrac{\vartheta}{2} - \tfrac{\rho}{Q}\bigr)\), which lies in \((0, \vartheta/2)\) because \(0 < \rho/Q < \vartheta/2\). Then \(\tfrac{\vartheta}{2} - \delta = \tfrac12\bigl(\tfrac{\vartheta}{2} + \tfrac{\rho}{Q}\bigr)\) and therefore \(\bigl(\tfrac{\vartheta}{2} - \delta\bigr)Q = \tfrac12(M + \rho) > \tfrac12(\rho + \rho) = \rho\), using \(M > \rho\). ◻
The window is non-empty for every \(\vartheta\) with \(2/\vartheta < Q\), with no further constraint tying \(\delta\) to \(\varepsilon\). Had one insisted instead on the reparametrisation condition \((1+\varepsilon)(\vartheta/2 - \delta) < \vartheta/2\), the window would have been empty in the decisive regime \(\delta \to 0\), \(\vartheta \to \tfrac12\), and the scaling would have yielded no improvement. It is the weak room \((1+\varepsilon)\vartheta < 1\) of Definition 1(iii), together with the support split, that keeps that regime available.
Proposition 1 (Witness form of the enlarged criterion). Let \(k \ge 2\), let \(\mathcal{H} = \{h_1, \dotsc, h_k\}\) be admissible, let \(0 \le \varepsilon < 1\), let \(F \in \mathcal{F}_\varepsilon\) with \(I_k(F) \ne 0\), and let \(\vartheta\) be \((\varepsilon, Q)\)-certificate-admissible for some \(Q\) with \(\mathcal{Q}_\varepsilon(F) \ge Q\). Then \(\mathrm{DHL}[k, 2]\) holds: there are infinitely many \(n\) for which at least two of \(n + h_1, \dotsc, n + h_k\) are prime.
Uses: prop_s1, prop_s2, def_M_k, def_cert_adm, def_prof_eps, def_admissible
Proof. By clauses (i) and (ii) of Definition 1 we have \(0 < \vartheta < \tfrac12\) and \(2/\vartheta < Q \le \mathcal{Q}_\varepsilon(F)\), so Lemma 1 supplies \(\delta \in (0, \vartheta/2)\) and \(\rho \in (1,2)\) with \(\bigl(\tfrac{\vartheta}{2} - \delta\bigr)\mathcal{Q}_\varepsilon(F) > \rho\). Fix these and set \(R := N^{\vartheta/2-\delta}\), so \(\log R / \log N = \vartheta/2 - \delta\).
Clause (iii) gives \(1 + \varepsilon < 1/\vartheta\), so Propositions 1 and 1 both apply. Summing Proposition 1 over \(m \in \{1,\dotsc,k\}\) and using \(\sum_m S_2^{(m)}(\lambda) = S_2(\lambda)\) (Lemma 1) together with Definition 1, \[S_2(\lambda_\varepsilon) \;\ge\; \bigl(1+o(1)\bigr)\, \frac{\varphi(W)^k N (\log R)^{k}}{W^{k+1}}\, \bigl(\tfrac{\vartheta}{2} - \delta\bigr)\, I_k(F)\, \mathcal{Q}_\varepsilon(F) ,\] where the exchange of one power of \(\log R\) for the factor \(\vartheta/2-\delta\) uses \((\log R)^{k+1}/(\log N \cdot (\log R)^k) = \vartheta/2 - \delta\). Proposition 1 gives \(S_1(\lambda_\varepsilon) = (1+o(1))\,\varphi(W)^kN(\log R)^kW^{-k-1} I_k(F)\), using \(\int_{\mathcal{T}_\varepsilon}F^2 = I_k(F)\) for \(F \in \mathcal{F}_\varepsilon\). Subtracting \(\rho\) times the latter from the former, \[S_2(\lambda_\varepsilon) - \rho\, S_1(\lambda_\varepsilon) \;\ge\; \bigl(1+o(1)\bigr)\, \frac{\varphi(W)^k N (\log R)^k}{W^{k+1}}\, I_k(F) \Bigl[ \bigl(\tfrac{\vartheta}{2} - \delta\bigr) \mathcal{Q}_\varepsilon(F) - \rho \Bigr] .\] The bracket is a fixed positive number by the choice of \(\rho\), and the prefactor is positive since \(I_k(F) \ne 0\) forces \(I_k(F) > 0\). Hence \(S(\lambda_\varepsilon) = S_2(\lambda_\varepsilon) - \rho\,S_1(\lambda_\varepsilon) > 0\) for every sufficiently large \(N\).
Finally, admissibility of \(\mathcal{H}\) supplies for each \(N\) a residue \(v_0\) modulo \(W\) with \(\gcd(v_0 + h_i, W) = 1\) for all \(i\), so the sieve sums above are the ones attached to a legitimate residue class, and Lemma 1 converts eventual positivity of \(S\) into infinitely many \(n\) for which at least \(\lfloor \rho + 1\rfloor\) of the \(n + h_i\) are prime. Since \(\rho \in (1,2)\) we have \(\lfloor \rho+1\rfloor = 2\), which is \(\mathrm{DHL}[k,2]\). ◻
Proof uses: lem_positivity_window, prop_s1, prop_s2, lem_S2_decomp, lem_infinitely_many
Proposition 1 demands a smooth witness, whereas the available certificates are polynomials cut off at the boundary of \(\mathcal{T}_\varepsilon\), hence not smooth. The gap is closed by mollification, using the \(L^2\)-continuity of \(\mathcal{Q}_\varepsilon\).
Lemma 1 (Mollification of an \(L^2\) witness). Let \(0 \le \varepsilon < 1\), let \(F_0 \in L^2(\mathcal{T}_\varepsilon)\) with \(I_k(F_0) \ne 0\), and let \(c < \mathcal{Q}_\varepsilon(F_0)\). Then there exists \(\widetilde F \in \mathcal{F}_\varepsilon\) with \(I_k(\widetilde F) \ne 0\) and \(\mathcal{Q}_\varepsilon(\widetilde F) > c\).
Proof. Set \(\eta := \mathcal{Q}_\varepsilon(F_0) - c > 0\) and let \(\zeta > 0\) be the radius supplied by Lemma 1 for \(F_1 := F_0\) and this \(\eta\): every \(F_2 \in L^2(\mathcal{T}_\varepsilon)\) with \(\|F_2 - F_0\|_{L^2(\mathcal{T}_\varepsilon)} < \zeta\) has \(I_k(F_2) \ne 0\) and \(\mathcal{Q}_\varepsilon(F_2) > \mathcal{Q}_\varepsilon(F_0) - \eta = c\).
It therefore suffices to produce \(\widetilde F \in \mathcal{F}_\varepsilon\) within \(\zeta\) of \(F_0\) in \(L^2(\mathcal{T}_\varepsilon)\). Transport the problem to the standard simplex by the homothety of Lemma 1: the map \(F \mapsto (1+\varepsilon)^{k/2}\,G_\varepsilon F\) is an isometry of \(L^2(\mathcal{T}_\varepsilon)\) onto \(L^2(\mathcal{R}_k)\), by Lemma 1 applied to differences, and it carries \(\mathcal{F}_\varepsilon\) onto the class of smooth functions supported on \(\mathcal{R}_k\). On \(L^2(\mathcal{R}_k)\), Lemma 1 provides smooth approximants to any given element within any prescribed \(L^2\) distance, and Lemma 1 lets one multiply such an approximant by a smooth cutoff supported on \(\mathcal{R}_k\) without increasing the distance by more than an arbitrarily small amount; the resulting function is smooth and supported on \(\mathcal{R}_k\). Pulling back through the isometry gives the required \(\widetilde F \in \mathcal{F}_\varepsilon\). ◻
Lemma 1 (Bombieri–Vinogradov supplies a certificate-admissible level). Let \(0 \le \varepsilon \le 1\) and let \(Q > 4\). Then there is an \((\varepsilon, Q)\)-certificate-admissible \(\vartheta\); one may take \(\vartheta = \tfrac12\bigl(\tfrac{2}{Q} + \tfrac12\bigr)\).
Uses: def_cert_adm, thm_BV
Proof. Since \(Q > 4\) we have \(2/Q < \tfrac12\), so the interval \((2/Q, \tfrac12)\) is non-empty; let \(\vartheta\) be its midpoint. Then \(0 < 2/Q < \vartheta < \tfrac12\), which is clause (i). From \(\vartheta > 2/Q\) and \(\vartheta, Q > 0\) we get \(2/\vartheta < Q\), clause (ii). For clause (iii), \(\vartheta < \tfrac12\) gives \(1/\vartheta > 2 \ge 1+\varepsilon\) whenever \(\varepsilon \le 1\); the inequality is strict because \(\vartheta < \tfrac12\) strictly. Finally \(\vartheta < \tfrac12\), so the primes have level of distribution \(\vartheta\) by Bombieri–Vinogradov (Theorem 1). ◻
Proof uses: thm_BV
The proof of Lemma 1 chooses \(\vartheta\) as a function of \(Q\) alone, and the closer \(Q\) is to \(4\) the closer \(\vartheta\) must be pushed to \(\tfrac12\); this is the only quantitative demand the enlarged theory makes on a witness. A witness attaining \(\mathcal{Q}_\varepsilon(F_0) > 4.0043\) permits \(\vartheta = 0.4995\), for which \(2/\vartheta = 4.00400\ldots\) and \((1+\varepsilon)\vartheta < 1\) for every \(\varepsilon \le 1\).
Proposition 1 (The enlarged criterion). Let \(k \ge 2\), let \(0 \le \varepsilon < 1\), and let \(\mathcal{H} = \{h_1,\dotsc,h_k\}\) be admissible. Suppose there exists \(F_0 \in L^2(\mathcal{T}_\varepsilon)\) with \[4\, \|F_0\|^2_{L^2(\mathcal{T}_\varepsilon)} \;<\; \sum_{m=1}^{k} J^{(m)}_\varepsilon(F_0) ,\] equivalently \(\mathcal{Q}_\varepsilon(F_0) > 4\), which by Lemma 1 certifies \(M_{k,\varepsilon} > 4\). Then \(\mathrm{DHL}[k,2]\) holds.
Proof. As in the proof of Lemma 1, the hypothesis forces \(\|F_0\|_{L^2(\mathcal{T}_\varepsilon)} \ne 0\), i.e. \(I_k(F_0) \ne 0\), and dividing gives \(\mathcal{Q}_\varepsilon(F_0) > 4\). Choose \(Q\) with \(4 < Q < \mathcal{Q}_\varepsilon(F_0)\). By Lemma 1 there is an \((\varepsilon, Q)\)-certificate-admissible \(\vartheta\). By Lemma 1 applied with \(c := Q\) there is \(\widetilde F \in \mathcal{F}_\varepsilon\) with \(I_k(\widetilde F) \ne 0\) and \(\mathcal{Q}_\varepsilon(\widetilde F) > Q\). Then \(\vartheta\) is \((\varepsilon, Q)\)-certificate-admissible and \(\mathcal{Q}_\varepsilon(\widetilde F) \ge Q\), so Proposition 1 applied to \(\widetilde F\) and \(\mathcal{H}\) gives \(\mathrm{DHL}[k,2]\).
At \(\varepsilon = 0\) every step above is the corresponding step of the standard development: by Lemma 1 the hypothesis reads \(4\,I_k(F_0) < \sum_m J^{(m)}_k(F_0)\), which is the standard criterion \(M_k > 4\) witnessed by \(F_0\). ◻
Proof uses: lem_M_eps_gt_four, lem_theta_bv, lem_mollify, prop_witness, lem_eps_zero_recovers
Proposition 1 is the criterion of the development, stated once in \((k, \varepsilon)\). At \(\varepsilon = 0\) it is verbatim the standard criterion — the demand \(M_k > 4\), witnessed by an explicit competitor rather than by the supremum — and the proof specialises to the standard proof, with the support split of Definition 1 becoming trivial (\(\lambda_{2,m} = 0\)) and clause (iii) of Definition 1 becoming vacuous. For \(\varepsilon > 0\) the competitor is allowed to live on a strictly larger region, at the cost of a strictly smaller marginal slice and of clause (iii); whether that trade is profitable at a given \(k\) is a question about the competitor, settled by exhibiting one. The two cases settle it at \((k, \varepsilon) = (105, 0)\) and at \((k, \varepsilon) = (50, \tfrac1{25})\) respectively.
The numerical certificate
The variational input to the \(\varepsilon\)-enlarged sieve is the single inequality \(M_{k,\varepsilon} > 4\). At \(k = 50\) this inequality is not established by an asymptotic estimate; one exhibits an explicit test function and computes its Rayleigh quotient exactly. The objects below reduce the inequality to an exact, finite computation with rational numbers.
There are two forms of certificate. The abstract form (Definition 1) is the form the argument uses: a single element of \(L^2(\mathcal{T}_\varepsilon)\) the sum of whose enlarged marginal functionals exceeds four times its squared norm. The explicit form (Definition 1) is what is checked: a finite list of exponent data together with a finite list of rational coefficients, subject to one inequality between two explicitly computable rational numbers. Lemmas 1 and 1 evaluate the two sides as exact rational quadratic forms. The analytic content lies in using the resulting test function; the certificate itself is arithmetic.
The certificate basis
Notation 1 (Signature). A signature is a finite multiset \(\alpha\) of positive integers, written in non-increasing order as \(\alpha = (\alpha_1 \ge \alpha_2 \ge \dotsb \ge \alpha_j \ge 1)\); the empty signature \(j = 0\) is admitted. Its size is \(|\alpha| := \sum_{i=1}^j \alpha_i\), its parts are the values \(\alpha_1, \dotsc, \alpha_j\), and for \(r \ge 1\) the deletion \(\alpha \setminus r\) is the signature obtained by removing one copy of the part \(r\) from \(\alpha\). By convention \(\alpha \setminus r := \alpha\) when \(r\) is not a part of \(\alpha\); in particular \(\alpha \setminus 0 = \alpha\).
Definition 1 (Monomial-symmetric polynomial). Let \(k \ge 1\) and let \(\alpha\) be a signature (Notation 1). Write \[E_k(\alpha) \;:=\; \bigl\{\, A = (A_1, \dotsc, A_k) \in \mathbb{Z}_{\ge 0}^k \;:\; \text{the multiset of non-zero entries of $A$ equals $\alpha$} \,\bigr\},\] a finite set, empty unless \(\alpha\) has at most \(k\) parts. The monomial-symmetric polynomial of signature \(\alpha\) in \(k\) variables is \[m_\alpha(t) \;:=\; \sum_{A \in E_k(\alpha)} \prod_{i=1}^k t_i^{A_i} \;\in\; \mathbb{Z}[t_1, \dotsc, t_k].\] Each distinct way of distributing the parts of \(\alpha\) among the \(k\) variables contributes exactly one monomial, so \(m_\alpha\) is symmetric, homogeneous of degree \(|\alpha|\), and \(m_\emptyset = 1\).
Uses: not_signature
Definition 1 (Certificate basis element). Let \(k \ge 1\), let \(\varepsilon \in \mathbb{R}\), let \(a \in \mathbb{Z}_{\ge 0}\) and let \(\alpha\) be a signature. The associated certificate basis element is the symmetric polynomial \[b_{a,\alpha}(t) \;:=\; \bigl(1 + \varepsilon - P_{(1)}(t)\bigr)^{a}\, m_\alpha(t), \qquad P_{(1)}(t) = t_1 + \dotsb + t_k,\] with \(m_\alpha\) as in Definition 1. Its total degree is \(a + |\alpha|\).
Definition 1 (Parametric certificate polynomial). Let \(k \ge 1\), \(\varepsilon \in \mathbb{R}\), and let \(n \in \mathbb{N}\). Fix basis indices \((a_1, \alpha_1), \dotsc, (a_n, \alpha_n)\), each consisting of an integer \(a_i \ge 0\) and a signature \(\alpha_i\), and let \(c = (c_1, \dotsc, c_n) \in \mathbb{Q}^n\). The associated certificate polynomial is \[F_c \;:=\; \sum_{i=1}^{n} c_i\, b_{a_i, \alpha_i} \;\in\; \mathbb{Q}[t_1, \dotsc, t_k],\] with \(b_{a,\alpha}\) as in Definition 1. It is a symmetric polynomial, hence continuous, hence bounded on the compact set \(\mathcal{T}_\varepsilon\); we regard it also as an element of \(L^2(\mathcal{T}_\varepsilon)\), and where a value off \(\mathcal{T}_\varepsilon\) is needed we use the extension of \(F_c\) by zero.
Uses: def_eps_basis, def_simplex
The two forms of certificate
Definition 1 (Abstract certificate). Let \(k \ge 1\) and \(\varepsilon \in \mathbb{R}\). An abstract certificate for \((k, \varepsilon)\) is a function \(F \in L^2(\mathcal{T}_\varepsilon)\) such that \[4\,\|F\|_{L^2(\mathcal{T}_\varepsilon)}^2 \;<\; \sum_{m=1}^{k} J^{(m)}_\varepsilon(F),\] where \(J^{(m)}_\varepsilon\) is the enlarged marginal functional of Definition 1. Equivalently, since \(\|F\|_{L^2(\mathcal{T}_\varepsilon)}^2 = I_k(F)\) for \(F\) extended by zero off \(\mathcal{T}_\varepsilon\), an abstract certificate is an \(F\) with \(\mathcal{Q}_\varepsilon(F) > 4\).
Uses: def_simplex, def_J_k, def_I_k, def_M_k
Definition 1 (Factorial moment of two signatures). Let \(k \ge 1\) and let \(\alpha, \beta\) be signatures. Their factorial moment in \(k\) variables is \[\Phi_k(\alpha, \beta) \;:=\; \sum_{A \in E_k(\alpha)} \sum_{B \in E_k(\beta)} \prod_{i=1}^{k} (A_i + B_i)! \;\in\; \mathbb{Z}_{\ge 0},\] the exponent sets \(E_k(\cdot)\) being those of Definition 1. It is a finite sum of positive integers, and \(\Phi_k(\emptyset, \emptyset) = 1\).
Definition 1 (Explicit enlarged-simplex pairing). Let \(k \ge 1\), \(\varepsilon \in \mathbb{Q}\), let \(a_1, a_2 \in \mathbb{Z}_{\ge 0}\) and let \(\alpha_1, \alpha_2\) be signatures. Put \(e := a_1 + a_2\) and \(S := |\alpha_1| + |\alpha_2|\). The explicit enlarged-simplex pairing is the rational number \[\mathcal{I}^{\varepsilon}_k(a_1, \alpha_1; a_2, \alpha_2) \;:=\; (1 + \varepsilon)^{\,k + e + S}\;\frac{e!}{(k + e + S)!}\;\Phi_k(\alpha_1, \alpha_2),\] with \(\Phi_k\) as in Definition 1.
Uses: def_fac_moment, not_signature
Definition 1 (Radial factor). For \(q, e \in \mathbb{Z}_{\ge 0}\) and \(\varepsilon \in \mathbb{Q}\), set \[\varrho_q(\varepsilon, e) \;:=\; (1 - \varepsilon)^{q} \sum_{r = 0}^{e} \binom{e}{r} (2\varepsilon)^{e - r} (1 - \varepsilon)^{r}\, \frac{r!}{(r + q)!} \;\in\; \mathbb{Q}.\]
Definition 1 (Explicit shrunken-slice pairing). Let \(k \ge 1\), \(\varepsilon \in \mathbb{Q}\), let \(a_1, a_2 \in \mathbb{Z}_{\ge 0}\) and let \(\alpha_1, \alpha_2\) be signatures. For parts \(r_1, r_2 \ge 0\) write \[q(r_1, r_2) \;:=\; (k - 1) + \bigl(|\alpha_1| - r_1\bigr) + \bigl(|\alpha_2| - r_2\bigr), \qquad e(r_1, r_2) \;:=\; (a_1 + r_1 + 1) + (a_2 + r_2 + 1).\] The explicit shrunken-slice pairing is the rational number \[\mathcal{J}^{\varepsilon}_k(a_1, \alpha_1; a_2, \alpha_2) \;:=\; \sum_{r_1} \sum_{r_2} \frac{r_1!\,a_1!}{(a_1 + r_1 + 1)!}\, \frac{r_2!\,a_2!}{(a_2 + r_2 + 1)!}\, \varrho_{q(r_1, r_2)}\bigl(\varepsilon, e(r_1, r_2)\bigr)\, \Phi_{k-1}\bigl(\alpha_1 \setminus r_1,\ \alpha_2 \setminus r_2\bigr),\] where \(r_1\) runs over \(0\) together with the distinct parts of \(\alpha_1\), and \(r_2\) over \(0\) together with the distinct parts of \(\alpha_2\); the deletions \(\alpha \setminus r\) and the convention \(\alpha \setminus 0 = \alpha\) are those of Notation 1, \(\varrho\) is Definition 1 and \(\Phi\) is Definition 1.
Definition 1 (Explicit certificate). Let \(k \ge 1\) and \(\varepsilon \in \mathbb{Q}\). An explicit certificate for \((k, \varepsilon)\) consists of
an integer \(n \ge 0\);
exponents \(a_1, \dotsc, a_n \in \mathbb{Z}_{\ge 0}\) and signatures \(\alpha_1, \dotsc, \alpha_n\) (Notation 1);
coefficients \(c_1, \dotsc, c_n \in \mathbb{Q}\);
subject to the single inequality \[4 \sum_{i=1}^{n}\sum_{j=1}^{n} c_i c_j\, \mathcal{I}^{\varepsilon}_k(a_i, \alpha_i; a_j, \alpha_j) \;<\; k \sum_{i=1}^{n}\sum_{j=1}^{n} c_i c_j\, \mathcal{J}^{\varepsilon}_k(a_i, \alpha_i; a_j, \alpha_j),\] with \(\mathcal{I}^{\varepsilon}_k\) as in Definition 1 and \(\mathcal{J}^{\varepsilon}_k\) as in Definition 1. Both sides are rational numbers determined by the data by finitely many additions and multiplications.
From a certificate to the bound \(M_{k,\varepsilon} > 4\)
Lemma 1 (Closed form of the enlarged-simplex pairing). Let \(k \ge 1\), let \(\varepsilon \in \mathbb{Q}\) with \(1 + \varepsilon \ge 0\), let \(a_1, a_2 \in \mathbb{Z}_{\ge 0}\) and let \(\alpha_1, \alpha_2\) be signatures (Notation 1). Then \[\int_{\mathcal{T}_\varepsilon} b_{a_1, \alpha_1}(t)\, b_{a_2, \alpha_2}(t) \, dt \;=\; \mathcal{I}^{\varepsilon}_k(a_1, \alpha_1; a_2, \alpha_2),\] with \(b_{a,\alpha}\) as in Definition 1 and \(\mathcal{I}^{\varepsilon}_k\) as in Definition 1. In particular the pairing is rational.
Proof. Expanding both monomial-symmetric factors (Definition 1) and collecting the two slack powers, \[b_{a_1, \alpha_1}(t)\, b_{a_2, \alpha_2}(t) \;=\; \sum_{A \in E_k(\alpha_1)} \sum_{B \in E_k(\alpha_2)} \bigl(1 + \varepsilon - \textstyle\sum_i t_i\bigr)^{a_1 + a_2} \prod_{i=1}^{k} t_i^{\,A_i + B_i},\] a finite sum of continuous functions, so the integral may be taken term by term. Here \(A\) and \(B\) denote exponent vectors, as in Definition 1. Fix \(A, B\), put \(C_i := A_i + B_i\) and write \(e := a_1 + a_2\). Every \(A \in E_k(\alpha_1)\) satisfies \(\sum_i A_i = |\alpha_1|\) and likewise for \(B\), so \(\sum_i C_i = |\alpha_1| + |\alpha_2| =: S\). Substituting \(t = (1 + \varepsilon) u\), which maps \(\mathcal{R}_k\) onto \(\mathcal{T}_\varepsilon\) and has Jacobian \((1+\varepsilon)^k\), and using \(1 + \varepsilon - \sum_i t_i = (1+\varepsilon)(1 - \sum_i u_i)\), \[\int_{\mathcal{T}_\varepsilon} \bigl(1 + \varepsilon - \textstyle\sum_i t_i\bigr)^{e} \prod_i t_i^{C_i}\, dt \;=\; (1+\varepsilon)^{\,k + e + S} \int_{\mathcal{R}_k} \bigl(1 - \textstyle\sum_i u_i\bigr)^{e} \prod_i u_i^{C_i}\, du .\] Lemma 1 evaluates the remaining integral as \(e! \prod_i C_i! / (k + e + S)!\). Summing over \(A \in E_k(\alpha_1)\) and \(B \in E_k(\alpha_2)\) turns \(\prod_i (A_i + B_i)!\) into \(\Phi_k(\alpha_1, \alpha_2)\) (Definition 1) and produces exactly \(\mathcal{I}^{\varepsilon}_k(a_1, \alpha_1; a_2, \alpha_2)\). ◻
Proof uses: def_monomial_symmetric, def_fac_moment, def_simplex, lem_monomial_integration
Lemma 1 (Closed form of the shrunken-slice pairing). Let \(k \ge 1\), let \(\varepsilon \in \mathbb{Q}\) with \(0 \le \varepsilon \le 1\), let \(m \in \{1, \dotsc, k\}\), let \(a_1, a_2 \in \mathbb{Z}_{\ge 0}\) and let \(\alpha_1, \alpha_2\) be signatures (Notation 1). Extending the basis elements by zero off \(\mathcal{T}_\varepsilon\), \[\int_{\mathcal{S}_\varepsilon} \Bigl(\int_0^\infty b_{a_1, \alpha_1}(t)\, dt_m\Bigr) \Bigl(\int_0^\infty b_{a_2, \alpha_2}(t)\, dt_m\Bigr)\, dt_{\hat m} \;=\; \mathcal{J}^{\varepsilon}_k(a_1, \alpha_1; a_2, \alpha_2),\] with \(b_{a,\alpha}\) as in Definition 1 and \(\mathcal{J}^{\varepsilon}_k\) as in Definition 1. In particular the value is rational and independent of \(m\).
Proof. Write \(\hat t = (t_i)_{i \ne m}\) for a point of \(\mathcal{S}_\varepsilon\) and set \(c(\hat t) := 1 + \varepsilon - \sum_{i \ne m} t_i\), the available length in the isolated coordinate. Since \(\sum_{i \ne m} t_i \le 1 - \varepsilon\) on \(\mathcal{S}_\varepsilon\), we have \(c(\hat t) \ge 2\varepsilon \ge 0\).
Step 1: the marginal of one basis element. A point with \(m\)-th coordinate \(s\) and remaining coordinates \(\hat t\) lies in \(\mathcal{T}_\varepsilon\) if and only if \(0 \le s \le c(\hat t)\), so the zero-extension makes \(\int_0^\infty b_{a,\alpha}\, dt_m = \int_0^{c(\hat t)} b_{a,\alpha}\, dt_m\). Expanding \(m_\alpha\) (Definition 1) and separating the isolated coordinate, \[b_{a,\alpha}\bigl(s; \hat t\bigr) \;=\; \sum_{A \in E_k(\alpha)} \bigl(c(\hat t) - s\bigr)^{a}\, s^{A_m} \prod_{i \ne m} t_i^{A_i} .\] By Lemma 1, after the substitution \(s = c(\hat t)\,v\), \[\int_0^{c} (c - s)^{a} s^{r}\, ds \;=\; \frac{a!\, r!}{(a + r + 1)!}\; c^{\,a + r + 1} \qquad (c \ge 0,\ a, r \in \mathbb{Z}_{\ge 0}),\] so that \[\Bigl(\int_0^\infty b_{a,\alpha}\, dt_m\Bigr)(\hat t) \;=\; \sum_{A \in E_k(\alpha)} \frac{a!\, A_m!}{(a + A_m + 1)!}\; c(\hat t)^{\,a + A_m + 1} \prod_{i \ne m} t_i^{A_i}.\]
Step 2: a shifted monomial integral over the shrunken slice. For \(e \in \mathbb{Z}_{\ge 0}\) and exponents \((C_i)_{i \ne m}\) we claim \[\int_{\mathcal{S}_\varepsilon} c(\hat t)^{\,e} \prod_{i \ne m} t_i^{C_i}\, dt_{\hat m} \;=\; \varrho_{q}(\varepsilon, e) \prod_{i \ne m} C_i!, \qquad q := (k-1) + \sum_{i \ne m} C_i,\] with \(\varrho\) as in Definition 1. Indeed \(c(\hat t) = 2\varepsilon + \bigl((1 - \varepsilon) - \sum_{i \ne m} t_i\bigr)\), so the binomial theorem gives \(c(\hat t)^e = \sum_{r=0}^{e} \binom{e}{r} (2\varepsilon)^{e-r} \bigl((1-\varepsilon) - \sum_{i \ne m} t_i\bigr)^{r}\). Each summand is integrated by the substitution \(\hat t = (1 - \varepsilon) u\), which maps \(\mathcal{R}_{k-1}\) onto \(\mathcal{S}_\varepsilon\) (here \(0 \le \varepsilon \le 1\) is used), followed by Lemma 1; this contributes \((1-\varepsilon)^{\,q + r}\, r! \prod_i C_i! / (r + q)!\), and summing over \(r\) gives the claim.
Step 3: assembly. Multiply the two marginals furnished by Step 1 and integrate over \(\mathcal{S}_\varepsilon\). The result is the double sum over \(A \in E_k(\alpha_1)\), \(B \in E_k(\alpha_2)\) of \[\frac{a_1!\,A_m!}{(a_1 + A_m + 1)!}\,\frac{a_2!\,B_m!}{(a_2 + B_m + 1)!}\, \int_{\mathcal{S}_\varepsilon} c(\hat t)^{\,(a_1 + A_m + 1) + (a_2 + B_m + 1)} \prod_{i \ne m} t_i^{\,A_i + B_i} \, dt_{\hat m},\] which Step 2 evaluates, since \(\sum_{i \ne m}(A_i + B_i) = (|\alpha_1| - A_m) + (|\alpha_2| - B_m)\). Finally group the double sum by the two isolated exponents \(r_1 := A_m\) and \(r_2 := B_m\): deleting the \(m\)-th coordinate is a bijection \[E_k(\alpha) \;\longleftrightarrow\; \coprod_{r} \{r\} \times E_{k-1}(\alpha \setminus r),\] where \(r\) runs over \(0\) together with the distinct parts of \(\alpha\) (Notation 1). Under this bijection the residual factorial products \(\prod_{i \ne m}(A_i + B_i)!\) sum to \(\Phi_{k-1}(\alpha_1 \setminus r_1, \alpha_2 \setminus r_2)\) (Definition 1), and the remaining factors are exactly those of Definition 1. No step depends on which coordinate was isolated, so the value is independent of \(m\). ◻
Lemma 1 (\(I_k(F_c)\) as a rational quadratic form). Let \(k \ge 1\), let \(\varepsilon \in \mathbb{Q}\) with \(1 + \varepsilon \ge 0\), and let \(F_c\) be the certificate polynomial attached to basis indices \((a_i, \alpha_i)_{i \le n}\) and coefficients \(c \in \mathbb{Q}^n\) (Definition 1). Then \[I_k(F_c) \;=\; \|F_c\|^2_{L^2(\mathcal{T}_\varepsilon)} \;=\; \sum_{i=1}^{n} \sum_{j=1}^{n} c_i c_j\, \mathcal{I}^{\varepsilon}_k(a_i, \alpha_i; a_j, \alpha_j) \;\in\; \mathbb{Q},\] where \(F_c\) is extended by zero off \(\mathcal{T}_\varepsilon\) in the definition of \(I_k\).
Proof. \(F_c\) is continuous, hence square-integrable on the compact set \(\mathcal{T}_\varepsilon\), and \(\|F_c\|^2_{L^2(\mathcal{T}_\varepsilon)} = \int_{\mathcal{T}_\varepsilon} F_c^2\). Expanding the square of the finite sum, \[F_c^2 \;=\; \sum_{i=1}^n \sum_{j=1}^n c_i c_j\, b_{a_i, \alpha_i}\, b_{a_j, \alpha_j},\] and integrating term by term (a finite sum of continuous functions on a compact set), Lemma 1 evaluates each \(\int_{\mathcal{T}_\varepsilon} b_{a_i,\alpha_i} b_{a_j,\alpha_j}\) as \(\mathcal{I}^{\varepsilon}_k(a_i, \alpha_i; a_j, \alpha_j)\). Rationality follows since each pairing and each \(c_i\) is rational. ◻
Proof uses: def_eps_basis, lem_eps_basis_pairing
Lemma 1 (\(\sum_m J^{(m)}_\varepsilon(F_c)\) as a rational quadratic form). Let \(k \ge 1\), let \(\varepsilon \in \mathbb{Q}\) with \(0 \le \varepsilon \le 1\), and let \(F_c\) be the certificate polynomial attached to basis indices \((a_i, \alpha_i)_{i \le n}\) and coefficients \(c \in \mathbb{Q}^n\) (Definition 1), extended by zero off \(\mathcal{T}_\varepsilon\). Then \[\sum_{m=1}^{k} J^{(m)}_\varepsilon(F_c) \;=\; k \sum_{i=1}^{n} \sum_{j=1}^{n} c_i c_j\, \mathcal{J}^{\varepsilon}_k(a_i, \alpha_i; a_j, \alpha_j) \;\in\; \mathbb{Q}.\]
Proof. Fix \(m\). The marginal \(\hat t \mapsto \int_0^\infty F_c\, dt_m\), taken for the extension of \(F_c\) by zero off \(\mathcal{T}_\varepsilon\), is linear in the coefficient vector: \[\int_0^\infty F_c \, dt_m \;=\; \sum_{i=1}^{n} c_i \int_0^\infty b_{a_i, \alpha_i}\, dt_m .\] Squaring and integrating over \(\mathcal{S}_\varepsilon\), and interchanging the finite sums with the integral, \[J^{(m)}_\varepsilon(F_c) \;=\; \sum_{i=1}^n \sum_{j=1}^n c_i c_j \int_{\mathcal{S}_\varepsilon} \Bigl(\int_0^\infty b_{a_i, \alpha_i} dt_m\Bigr)\Bigl(\int_0^\infty b_{a_j, \alpha_j} dt_m\Bigr) d\hat t \;=\; \sum_{i=1}^n \sum_{j=1}^n c_i c_j\, \mathcal{J}^{\varepsilon}_k(a_i, \alpha_i; a_j, \alpha_j)\] by Lemma 1. The right-hand side does not depend on \(m\), so summing over the \(k\) coordinates multiplies it by \(k\). ◻
Proof uses: def_eps_basis, lem_eps_basis_marginal_pairing
Existence of a certificate at \(k = 50\)
Proposition 1 (Existence of an explicit certificate for \(k = 50\)). There exist a rational number \(\varepsilon\) with \(0 \le \varepsilon \le 1\) and an explicit certificate for \((50, \varepsilon)\) in the sense of Definition 1.
Uses: def_explicit_certificate
Proof. Take \(\varepsilon = 1/25\), so that \(0 \le \varepsilon \le 1\). Let the index set be \[\mathcal{B} \;:=\; \bigl\{\, (a, \alpha) \;:\; a \in \mathbb{Z}_{\ge 0},\ \alpha \text{ a signature all of whose parts are even and $\ge 2$},\ a + |\alpha| \le 27 \,\bigr\},\] a finite set; enumerate it in any fixed order as \((a_1, \alpha_1), \dotsc, (a_n, \alpha_n)\) with \(n = |\mathcal{B}| = 2526\). Definitions 1 and 1 attach to this enumeration two explicit rational symmetric matrices, \[(A_I)_{ij} := \mathcal{I}^{\varepsilon}_{50}(a_i, \alpha_i; a_j, \alpha_j), \qquad (A_J)_{ij} := 50\,\mathcal{J}^{\varepsilon}_{50}(a_i, \alpha_i; a_j, \alpha_j),\] each entry being a finite sum of products of factorials, binomial coefficients and powers of \(\varepsilon\). What has to be produced is a vector \(c \in \mathbb{Q}^n\) with \(4\, c^{\top} A_I\, c < c^{\top} A_J\, c\); once \(c\) is written down, this is a comparison of two explicitly given rational numbers and is decided by exact integer arithmetic after clearing denominators.
Such a vector is found as follows. By Lemma 1, \(c^{\top} A_I\, c = \int_{\mathcal{T}_\varepsilon} F_c^2 \ge 0\), and it vanishes only for \(c = 0\) because the polynomials \(b_{a,\alpha}\), \((a,\alpha) \in \mathcal{B}\), are linearly independent; so \(A_I\) is positive definite and the generalised eigenvalue problem \(A_J v = \lambda A_I v\) has real eigenvalues, the largest being \(\lambda_{\max} = \sup_{v \ne 0} (v^{\top} A_J v)/(v^{\top} A_I v)\). Computing a top eigenvector to sufficient precision and rounding it coordinatewise to a rational vector over a common denominator yields \(c\). None of this is part of the verification: the precision analysis, the eigenvector and the rounding only serve to produce a candidate \(c\), and once \(c\) is fixed the sole assertion to be checked is the displayed rational inequality. Carrying this out at \((k, \varepsilon, d) = (50, 1/25, 27)\) gives a \(c\) with \[\frac{c^{\top} A_J\, c}{c^{\top} A_I\, c} \;>\; 4.0043 \;>\; 4,\] so the inequality of Definition 1 holds. The degree bound cannot be relaxed by much: the supremum of the quotient over the span of the degree-\(d\) basis increases with \(d\), and at \((k, \varepsilon) = (50, 1/25)\) it is still below \(4\) for \(d \le 24\), its value there being \(3.9995\ldots\); the exhibited \(c\) lives at \(d = 27\), where the exact quotient is \(4.00430\ldots\). The restriction to signatures with even parts is likewise not cosmetic: it is what the test functions used here look like, and dropping the multi-part signatures (allowing only \(\alpha\) with at most two parts, say) lowers the attainable quotient below \(4\).
It is equivalent, and closer to the classical formulation, to search in the power-sum monomials \(P_\alpha = \prod_j P_{\alpha_j}\) with \(P_r = \sum_{i} t_i^r\) instead of the monomial-symmetric polynomials \(m_\alpha\): expanding a product of power sums groups the coordinates into blocks, so \(P_\alpha\) is a non-negative integer combination of the \(m_\mu\) with \(\mu\) obtained by merging parts of \(\alpha\); such \(\mu\) again have even parts \(\ge 2\) and \(|\mu| = |\alpha|\). Hence \((1 + \varepsilon - P_{(1)})^a P_\alpha\) lies in the span of \(\{b_{a', \alpha'} : (a', \alpha') \in \mathcal{B}\}\) for every \((a, \alpha) \in \mathcal{B}\), and a rational coefficient vector in the power-sum family transforms, by a finite integer matrix, into one in the family \(\mathcal{B}\). ◻
Proposition 1 gives an \(\varepsilon \in \mathbb{Q}\cap [0,1]\) and an explicit polynomial \(F \in L^2(\mathcal{T}_\varepsilon)\) whose enlarged Rayleigh quotient exceeds \(4\); the polynomial is then smoothed and used in the enlarged sieve estimates.
The admissible 50-tuple
Definition 1 (The tuple \(\mathcal{H}_{50}\)). Let \[\begin{aligned} \mathcal{H}_{50}\ :=\ \{\, & 0,\, 4,\, 6,\, 16,\, 30,\, 34,\, 36,\, 46,\, 48,\, 58,\, 60,\, 64,\, 70,\, 78,\, 84,\, 88,\, 90,\\ & 94,\, 100,\, 106,\, 108,\, 114,\, 118,\, 126,\, 130,\, 136,\, 144,\, 148,\, 150,\, 156,\, 160,\\ & 168,\, 174,\, 178,\, 184,\, 190,\, 196,\, 198,\, 204,\, 210,\, 214,\, 216,\, 220,\, 226,\, 228,\\ & 234,\, 238,\, 240,\, 244,\, 246\,\}\ \subset\ \mathbb{Z}_{\ge 0}, \end{aligned}\] a set of \(50\) distinct non-negative integers with \(\min\mathcal{H}_{50}=0\) and \(\max\mathcal{H}_{50}=246\). It is the narrowest known admissible \(50\)-tuple, normalised so that its least element is \(0\); see the Engelsma–Sutherland tables of admissible tuples, https://math.mit.edu/~primegaps/tuples/admissible_50_246.txt.
Lemma 1 (Diameter of \(\mathcal{H}_{50}\)). \(\operatorname{diam}(\mathcal{H}_{50})=246\).
Proof. The elements of \(\mathcal{H}_{50}\) are listed in increasing order in Definition 1, so \(\min\mathcal{H}_{50}=0\) and \(\max\mathcal{H}_{50}=246\). By Definition 1, \(\operatorname{diam}(\mathcal{H}_{50})=\max\mathcal{H}_{50}-\min\mathcal{H}_{50}=246-0=246\). ◻
Lemma 1 (Admissibility of \(\mathcal{H}_{50}\)). \(\mathcal{H}_{50}\) is admissible.
Uses: def_h50, def_admissible
Proof. By Lemma 1 it suffices to show that \(\#\{h\bmod p:h\in\mathcal{H}_{50}\}<p\) for every prime \(p\).
For \(p>50\) this is automatic: the reduction map \(\mathcal{H}_{50}\to\mathbb{Z}/p\mathbb{Z}\) has image of cardinality at most \(\#\mathcal{H}_{50}=50<p\). Thus only the fifteen primes \[p\in\{2,3,5,7,11,13,17,19,23,29,31,37,41,43,47\}\] require checking, and for each of them the reduction of the explicit list of Definition 1 is a finite computation. It omits a residue class in every case; for instance every element of \(\mathcal{H}_{50}\) is even, so the image modulo \(2\) is \(\{0\}\), and modulo \(3\) the image is \(\{0,1\}\), missing \(2\). ◻
Proof uses: def_h50, lem_admissible_cardinality_char
Assembly
The two bounds below are two cases of one argument. Its variational input is a lower bound for the variational supremum; its arithmetic input is an admissible tuple of the corresponding size. Taking the standard simplex, \(k=105\), the bound \(M_{105}>4\), and the tuple \(\mathcal{H}_{105}\) of diameter \(600\) gives Theorem 4. Taking the \(\varepsilon\)-enlarged simplex, \(k=50\), a numerical certificate of \(\mathcal{Q}_\varepsilon>4\), and the tuple \(\mathcal{H}_{50}\) of diameter \(246\) gives Theorem 8. In the second case the variational input is a finite, exact verification (Definition 1, Theorem 7).
Definition 1 (A certificate at \(k=50\)). A \(50\)-certificate is a datum consisting of a rational number \(\varepsilon\) with \(0\le\varepsilon\le1\), finitely many elements \(b_1,\dotsc,b_n\) of the certificate basis of Definition 1 in the \(50\) variables \(t_1,\dotsc,t_{50}\), and a rational vector \(a=(a_1,\dotsc,a_n)\in\mathbb{Q}^n\), such that the polynomial \(F_a=\sum_{i=1}^na_ib_i\) of Definition 1 satisfies the strict inequality \[4\,I_{50}(F_a)\ <\ \sum_{\ell=1}^{50}J^{(\ell)}_\varepsilon(F_a).\]
Uses: def_eps_basis, def_cert_polynomial, def_I_k, def_J_k
Both sides of the displayed inequality are rational numbers determined by the finite datum \((\varepsilon,(b_i)_{i\le n},a)\) through the closed-form pairings of Lemmas 1 and 1, so the condition defining a \(50\)-certificate is a single exact comparison of two rationals. The analytic argument uses only that a \(50\)-certificate exists, not the particular \(\varepsilon\), basis, or coefficient vector.
Theorem 5 (\(\mathrm{DHL}[50,2]\) from a certificate). Suppose the primes have level of distribution \(\vartheta\) for every \(\vartheta<\tfrac12\) (Theorem 1), and suppose a \(50\)-certificate exists (Definition 1). Then \(\mathrm{DHL}[50,2]\) holds: for every admissible \(50\)-tuple \(\mathcal{H}=\{h_1<\dotsb<h_{50}\}\) there are infinitely many \(n\in\mathbb{N}\) such that at least two of \(n+h_1,\dotsc,n+h_{50}\) are prime.
Uses: def_cert_50, def_admissible, thm_BV
Proof. Fix a \(50\)-certificate \((\varepsilon,(b_i)_{i\le n},a)\) and set \(F_0:=F_a\), so that \[\label{eq_cert_ineq} 4\,I_{50}(F_0)\ <\ \sum_{\ell=1}^{50}J^{(\ell)}_\varepsilon(F_0).\] The functional \(I_{50}\) is an integral of a square, so \(I_{50}(F_0)\ge0\); and \(I_{50}(F_0)=0\) would force \(F_0=0\) almost everywhere, hence \(J^{(\ell)}_\varepsilon(F_0)=0\) for every \(\ell\), contradicting [eq_cert_ineq]. Therefore \(I_{50}(F_0)>0\) and the witness ratio \(Q:=\mathcal{Q}_\varepsilon(F_0)\) of Definition 1 is defined and satisfies \(Q>4\) by [eq_cert_ineq].
Apply Lemma 1 with this \(\varepsilon\in[0,1]\) and this \(Q>4\): it produces a level \(\vartheta<\tfrac12\) with \(2/\vartheta<Q\) and \(1+\varepsilon<1/\vartheta\), at which the primes have level of distribution \(\vartheta\). Note \(\vartheta>0\), since \(1/\vartheta>1+\varepsilon\ge1>0\).
Apply Lemma 1 at this \(\vartheta\): since \(\mathcal{Q}_\varepsilon(F_0)>2/\vartheta\) and \(1+\varepsilon<1/\vartheta\), it produces a smooth \(F\) supported on the enlarged simplex \(\mathcal{T}_\varepsilon\) and a margin \(\delta\in(0,\vartheta/2)\) with \[\int_{\mathcal{T}_\varepsilon}F^2\ <\ \Bigl(\frac{\vartheta}{2}-\delta\Bigr)\sum_{m=1}^{50}J^{(m)}_\varepsilon(F).\]
Now let \(\mathcal{H}=\{h_1<\dotsb<h_{50}\}\) be an arbitrary admissible \(50\)-tuple. The data \(\varepsilon\), \(F\), \(\vartheta\), \(\delta\) satisfy every hypothesis of Proposition 1 at \(k=50\) with \(\rho=1\): \(\varepsilon\in[0,1]\); \(F\) is smooth and supported on \(\mathcal{T}_\varepsilon\); \(\vartheta\in(0,1)\), \(\vartheta<\tfrac12\), and the primes have level of distribution \(\vartheta\); \(\delta\in(0,\vartheta/2)\); \(1+\varepsilon<1/\vartheta\); and the displayed inequality is exactly the witness inequality with \(\rho=1\). Proposition 1 therefore gives infinitely many \(n\) with at least two of \(n+h_1,\dotsc,n+h_{50}\) prime. As \(\mathcal{H}\) was arbitrary, this is \(\mathrm{DHL}[50,2]\). ◻
Proof uses: def_M_k, lem_theta_bv, lem_mollify, prop_witness
Theorem 6 (\(H_1\le246\), conditional on a certificate). Suppose the primes have level of distribution \(\vartheta\) for every \(\vartheta<\tfrac12\) (Theorem 1), and suppose a \(50\)-certificate exists (Definition 1). Then \[\liminf_{m\to\infty}\bigl(p_{m+1}-p_m\bigr)\ \le\ 246 .\]
Uses: thm_BV, def_cert_50
Proof. The tuple \(\mathcal{H}_{50}\) is admissible (Lemma 1) and has \(\operatorname{diam}(\mathcal{H}_{50})=246\) (Lemma 1). Both hypotheses of Theorem 5 hold, so \(\mathrm{DHL}[50,2]\) holds, and applying it to \(\mathcal{H}_{50}\) shows that \[\mathcal{N}\ :=\ \{n\in\mathbb{N}:\ \text{at least two of the shifts } n+h,\ h\in\mathcal{H}_{50},\ \text{are prime}\}\] is infinite.
For \(n\in\mathcal{N}\) every shift \(n+h\) with \(h\in\mathcal{H}_{50}\) lies in \([\,n+\min\mathcal{H}_{50},\ n+\max\mathcal{H}_{50}\,]\), an interval of length \(\operatorname{diam}(\mathcal{H}_{50})=246\). Hence \(\mathcal{N}':=\{n+\min\mathcal{H}_{50}:n\in\mathcal{N}\}\) is an infinite set of natural numbers such that \([n',n'+246]\) contains at least two distinct primes for every \(n'\in\mathcal{N}'\).
Lemma 1 applied to \(\mathcal{N}'\) with \(r=2\) and \(L=246\) gives the conclusion. ◻
Proof uses: thm_dhl_50, lem_h50_adm, lem_h50_diam, lem_many_primes_implies_small_gaps
Theorem 7 (Existence of a \(50\)-certificate). A \(50\)-certificate exists (Definition 1).
Uses: def_cert_50
Proof. Take \(\varepsilon=\tfrac1{25}\), which is rational and lies in \([0,1]\), and take for \(b_1,\dotsc,b_n\) the certificate basis of Definition 1 at \(k=50\), \(\varepsilon=\tfrac1{25}\) and total-degree bound \(27\); here \(n=2526\). Proposition 1 supplies a vector \(a^\ast\in\mathbb{Q}^n\) with \(I_{50}(F_{a^\ast})>0\) and \(\mathcal{Q}_{1/25}(F_{a^\ast})>4.0043\). By Definition 1 this reads \[\sum_{\ell=1}^{50}J^{(\ell)}_{1/25}(F_{a^\ast})\ >\ 4.0043\; I_{50}(F_{a^\ast})\ >\ 4\,I_{50}(F_{a^\ast}),\] the last step because \(I_{50}(F_{a^\ast})>0\). Thus \((\tfrac1{25},(b_i)_{i\le n},a^\ast)\) is a \(50\)-certificate.
The verification is finite and exact. By Lemmas 1 and 1 the two sides are the rational quadratic forms \((a^\ast)^\top A_I\,a^\ast\) and \((a^\ast)^\top A_J\,a^\ast\) in the pairing matrices of Definitions 1 and 1, whose entries are given in closed form by Lemmas 1 and 1. With \(a^\ast\) taken over a common denominator, clearing denominators turns the required inequality into a single comparison of two integers, namely \(10^4\cdot(a^\ast)^\top A_J\,a^\ast>4\,0043\cdot(a^\ast)^\top A_I\,a^\ast\); its exact value is \(\mathcal{Q}_{1/25}(F_{a^\ast})=4.00430625\dotso\), which clears \(4.0043\), and a fortiori \(4\), strictly. No analytic input enters at this point: the statement is a finite assertion about rational numbers. ◻
Theorem 8 (\(H_1\le246\)). Suppose the primes have level of distribution \(\vartheta\) for every \(\vartheta<\tfrac12\) (Theorem 1). Then \[\liminf_{m\to\infty}\bigl(p_{m+1}-p_m\bigr)\ \le\ 246 .\]
Uses: thm_BV
Proof. Theorem 7 provides a \(50\)-certificate, which is the second hypothesis of Theorem 6; the first is the present hypothesis. Theorem 6 then gives the conclusion. ◻
Proof uses: thm_246_cert, thm_cert_50
99 J. Maynard, Small gaps between primes, Ann. of Math. (2) 181 (2015), no. 1, 383–413. D. H. J. Polymath, Variants of the Selberg sieve, and bounded intervals containing many primes, Res. Math. Sci. 1 (2014), Art. 12, arXiv:1407.4897.