# Problem Frontier Baseline Audit — 20 July 2026

Audit completed: `2026-07-21`

Status: `BASELINE REPRODUCIBLE / WORKING COUNT CORRECTED / TREND NOT READY / ARTICLE HOLD`

## Decision in one paragraph

The workpack has earned a reproducible methods baseline, but it has not earned
the article's implied depletion story. The frozen 20 July tracker correctly
preserves the nine AlphaProof Nexus Erdős rows initially presented as new
full-resolution events. Individual audit subsequently found that one of those
rows, Erdős 846, was already implied by published pre-2026 work. The corrected
working count is therefore eight strict-core resolutions from 353 attempted
formalized Erdős problems, not nine. This is a meaningful AI-mathematics result
and an equally meaningful measurement correction. It is not evidence that AI
is exhausting mathematics, that the resolution rate is accelerating or that a
human problem frontier is shrinking faster than it is replenished.

## What this audit covers

This audit tests five layers separately:

1. whether the frozen baseline can be reproduced from immutable inputs;
2. whether the event rows correspond to the original mathematical questions;
3. whether the proof artifacts and prose artifacts support the recorded result;
4. whether each result was genuinely open immediately before the contribution;
   and
5. whether the resulting counts support any rate, trend or depletion claim.

It does not independently re-prove all eight retained theorems from first
principles. Mechanical verification, original-source comparison, public
chronology, correction history, prose alignment and available external
mathematical response are recorded row by row. Several retained results still
require a broader literature-priority review before a public article could use
them without qualification.

## Baseline identity and custody

The accepted frozen object is:

- snapshot ID: `PFT-2026-07-20`;
- snapshot date: `2026-07-20`;
- status: `baseline_with_open_audit_questions`;
- next scheduled comparison: `2026-10-20`;
- snapshot SHA-256:
  `25b3e006bf9535e84ea4d3e97c037160506aa88bb124d1903f30007fda1cb0f9`.

The tracker freezes its own copies of the event table, universe table, source
receipts, registry observations, Erdős artifact inventory and OEIS inventory.
Running the tracker verifier against those copied inputs reproduces the
snapshot exactly. The old snapshot has not been regenerated from the corrected
working data.

That distinction is essential:

- **frozen baseline:** what the seed and artifact inventory said on 20 July;
- **audit correction:** what row-level review established on 21 July; and
- **future comparison:** a later accepted snapshot that must carry the
  correction forward as a reclassification rather than rewriting history.

The current correction is not a second longitudinal observation. It is an
audit layer over the first observation.

## Frozen result versus corrected working result

| Measure | Frozen 20 July snapshot | Corrected working table | Reading |
|---|---:|---:|---|
| Erdős result artifacts | 9 | 9 | All nine Lean/PDF pairs exist. |
| strict-core full resolutions | 9 | 8 | Erdős 846 fails the novelty/open-status gate. |
| attempted Erdős records | 353 | 353 | Fixed experimental denominator reproduced. |
| strict-core cohort ratio | 2.5496% | 2.2663% | Experimental yield, not global mathematical depletion. |
| median known problem age | 32 years | 32 years | Problem 26 has no defensible source year and is omitted from the median. |
| core quarters represented | 1 | 2 | Corrected public dates produce Q1 and Q2, still not an adequate time series. |
| trend ready | no | no | No comparable multi-period series or replenishment measure exists. |

The supplied seed originally assigned all nine rows the paper publication date
in Q2. The row audit replaced that convenience date with each first located
public proof or result date. The corrected table contains two retained events
in `2026-Q1` and six in `2026-Q2`. These are releases from one short research
campaign by one system, under one selected formalized cohort. Two adjacent
quarters do not support constant-versus-linear-versus-exponential model
comparison.

## Nine-row Erdős audit

| Event | Counted result | Statement and artifact finding | Novelty/open-status finding | Core decision |
|---|---|---|---|---|
| E-003 | 12(i) | Original subproblem passes; official CI passes; PDF uses a different construction. | Curated support; broader priority review pending. | include |
| E-004 | 12(ii) | Original subproblem passes; official CI passes; PDF overstates its quantifier order. | Curated support; broader priority review pending. | include |
| E-005 | 125 lower-density variant | Corrected intended variant passes after the first natural-density target failed to settle it; Lean and PDF align. | Named external response; upper-density question remains open. | include |
| E-006 | 138 difference variant | Exact named variant passes and proves the stronger pointwise bound; root problem remains open. | Named external interpretation; broader priority review pending. | include |
| E-007 | 152 | Finite uniform statement passes with an explicit quadratic bound; Lean and PDF align. | A public prior-art objection was mathematically withdrawn and the result reconfirmed. | include |
| E-011 | 26 thick-set weak-density variant | Formal variant passes; Lean strengthens the shift domain; PDF uses a related but non-identical density argument. | Direct source for the Tenenbaum attribution and problem year remains missing. | include |
| E-008 | 741(i) upper-density interpretation | Corrected upper-density theorem passes; first natural-density result is specification history; public artifact mapping points to the wrong Lean proof. | Upper version counted; lower-density interpretation remains open. | include |
| E-009 | 741(ii) | Original basis convention passes; released Lean and first-proof/PDF constructions differ; PDF has a minor omitted infinite-colour step. | Independent same-day model proof corroborates the result; broader priority review pending. | include |
| E-010 | 846 | Exact 1992 statement and official CI pass; PDF omits the formal proof's major non-collinearity case analysis. | Reiher–Rödl–Sales already implied the negative answer after specialization and generic projection. | exclude |

The detailed row reports are:

- [12(i)](ERDOS_ROW_AUDIT_12_I_2026-07-20.md)
- [12(ii)](ERDOS_ROW_AUDIT_12_II_2026-07-21.md)
- [125 lower density](ERDOS_ROW_AUDIT_125_LOWER_DENSITY_2026-07-21.md)
- [138 difference variant](ERDOS_ROW_AUDIT_138_DIFFERENCE_2026-07-21.md)
- [152](ERDOS_ROW_AUDIT_152_2026-07-21.md)
- [26 weak-density variant](ERDOS_ROW_AUDIT_26_WEAK_DENSITY_2026-07-21.md)
- [741(i) upper density](ERDOS_ROW_AUDIT_741_I_UPPER_DENSITY_2026-07-21.md)
- [741(ii)](ERDOS_ROW_AUDIT_741_II_2026-07-21.md)
- [846](ERDOS_ROW_AUDIT_846_2026-07-21.md)

## Why Erdős 846 changes the metric

Erdős 846 remains strong evidence about model capability. AlphaProof Nexus
independently generated a complete explicit construction, formalized the
hardest determinant case analysis and produced a proof accepted by official
repository CI. A separate internal OpenAI model found another explicit proof.

It is not a strict frontier resolution under this workpack's own gate. The
OpenAI paper records Vojtěch Rödl's observation that Theorem 1.7 of Reiher,
Rödl and Sales already implies the plane counterexample after the `k = 3`
specialization and generic projection. The relevant literature predates both
AI constructions. The living registry had not absorbed that implication, but
registry state cannot override mathematical priority.

The correct classification is therefore:

- `validation_status = formally_verified`;
- `novelty_status = known_in_literature_not_in_registry`; and
- `inclusion_status = exclude_not_open`.

This does not call the agent proof trivial, copied or mathematically invalid.
It says that “independently found a new proof” and “resolved a still-open
frontier problem” are different claims.

Primary references:

- [original 1992 problem](https://lematematiche.dmi.unict.it/index.php/lematematiche/article/view/587);
- [first public AlphaProof Nexus proof commit](https://github.com/google-deepmind/formal-conjectures/commit/2404258180688283e5141021c75464dc2acfb798);
- [independent OpenAI-model proof and Rödl observation](https://arxiv.org/abs/2602.21275v1); and
- [Reiher–Rödl–Sales result](https://arxiv.org/abs/2311.08556v2).

## Denominator audit

### U-003: AlphaProof Nexus Erdős attempted cohort

The `353` denominator reproduces from the pinned attempted-record file and is
the best current denominator in the pack. It is valid only for the selected
AlphaProof Nexus experiment cohort.

It is not:

- the number of all open Erdős problems;
- the number of all formalized open mathematical problems;
- a random or representative sample of mathematics; or
- a stock from which eight problems can simply be subtracted to estimate when
  mathematics will run out.

The publication-safe measure is “eight strict-core resolutions among 353
attempted formalized Erdős records in this experiment,” not “2.27% of
mathematics has been depleted.”

### U-004: AlphaProof Nexus OEIS attempted cohort

The denominator `492` reproduces from the pinned theorem mapping. The claimed
success numerator does not:

- paper claim: `44` successful conjectures;
- public Lean result files: `38`;
- represented unique OEIS sequence IDs: `37`; and
- unexplained claim-to-file gap: `6`.

Repository history, paper source, supplements and the public issue thread do
not close the gap. The `44/492` statement may be reported only as an author
claim. It cannot enter the audited event curve until the six records or the
counting rule are supplied. See
[the OEIS reconciliation](OEIS_RECONCILIATION_2026-07-20.md).

### U-005: living Erdős registry

At the pinned revision, the repository contains:

- `1,217` total problem records;
- `610` records in the completely-open state; and
- `663` records across the explicitly enumerated open-like components.

The supplied `627` completely-open note and encoded `670` eligible count are
not current under the pinned source. Because the registry is living and its
status classes differ conceptually, it cannot be silently substituted for the
fixed `353` experimental denominator.

### U-001 and U-002: Formal Conjectures

The paper reports `1,029` open research conjectures, but a problem-level frozen
membership table has not yet been built at the matching paper revision. The
living repository is pinned, but its current membership is not interchangeable
with the paper cohort. Neither is ready for a longitudinal depletion curve.

## Validation and contribution audit

All eight retained core rows are formally verified. That is a strong and
specific statement: the released proof assistant accepts the formal target at
the pinned repository state or corrected proof commit.

It does not establish by itself:

- that the formal target is the historically intended problem;
- that the result was absent from all prior literature;
- that the public prose is a complete proof;
- that no decisive human intervention occurred inside an unpublished run; or
- that the selected attempted cohort represents frontier difficulty.

The nine-row audit found examples of every relevant boundary:

- statement clarification changed the counted target for 125 and 741(i);
- variants had to be separated from composite root problems for 138 and 26;
- prose and formal constructions differed for 12(i), 12(ii), 741(ii) and 26;
- public artifact links were wrong for both 741 rows;
- a prose proof omitted a major lemma in 846; and
- literature priority removed a formally correct result from the strict core.

These are not reasons to discard formal mathematics. They are evidence that
formal verification and frontier accounting are different jobs.

## Trend audit

Trend readiness remains `false`.

The current table fails the workpack's minimum conditions:

- only eight individually audited strict-core events are included, below the
  minimum of twenty;
- the corrected public dates occupy only two adjacent quarters;
- all eight belong to one system and one experiment cohort;
- there is no accepted second tracker snapshot;
- no comparable replenishment series exists;
- unsuccessful attempts are known only as membership in an attempted cohort,
  not as comparable per-run histories;
- compute budgets and selection effects are unavailable; and
- the OEIS numerator remains unreconciled.

Fitting an exponential curve here would turn release scheduling and source
selection into a capability law. Constant, linear, exponential and step-change
models cannot be meaningfully compared with this observation window.

## Falsification and kill-condition result

The main measured claim currently fails its publication gate.

| Test | Baseline result |
|---|---|
| accelerating verified resolution rate | not testable |
| meaningful fixed-cohort depletion | only an experiment-specific `8/353`; no longitudinal decline |
| difficulty shift | not measured |
| human scaffolding sensitivity | incomplete; public run logs absent |
| novelty stability | failed for one of nine reported Erdős artifacts |
| statement stability | several corrections and variant boundaries found |
| validation attrition | measurable in principle, but no multi-period series |
| reporting bias | material; failures and OEIS gap are incompletely exposed |
| compute confound | not measurable from current evidence |
| replenishment | no accepted measure |
| human comprehension | not operationalized |

Three explicit kill conditions already apply to any empirical suggestion of
imminent depletion:

1. fewer than twenty independently validated individual full resolutions;
2. an observation window too short for trend comparison; and
3. no defensible replenishment measure.

The title may remain a question. The data may not answer it with “soon.”

## Claims the baseline can support

The current evidence supports these claims:

1. AI systems have produced formally verified solutions to genuine,
   human-formulated research problems.
2. AlphaProof Nexus has eight currently retained strict-core Erdős resolutions
   among 353 attempted formalized records after individual audit.
3. Formal verification does not settle statement fidelity, prose completeness,
   novelty or contribution provenance.
4. A living open-problem registry can lag published implications, as Erdős 846
   demonstrates.
5. The work of measuring machine mathematics is itself a source-audit problem,
   not merely a matter of counting green rows.

## Claims the baseline cannot support

The current evidence does not support:

- AI is running out of mathematical problems;
- AI is exhausting the mathematical frontier;
- the verified resolution rate is exponential or accelerating;
- eight of 353 represents all mathematics or all open Erdős problems;
- AlphaProof Nexus has an independently reproduced `44/492` OEIS result rate;
- every formally verified result was novel when produced;
- public artifact prose faithfully exposes every formal proof; or
- human problem creation is slower than AI problem resolution.

## Publication decision

Recommended state: `HOLD ARTICLE / METHODS NOTE CANDIDATE`.

The audit is publishable in principle as a transparent methods finding: one
reported nine-result line becomes eight under its own strict definition, the
OEIS numerator does not publicly reproduce, and no trend can yet be inferred.
That is a stronger Hedegreen Research contribution than forcing the current
evidence to confirm the title.

The full article remains blocked until Dennis reviews this audit and chooses
one of three paths:

1. publish a narrow methods note about how AI mathematical results should be
   counted;
2. continue the research program until a comparable second snapshot and
   replenishment measure exist; or
3. retain the workpack as an internal register and reject the depletion thesis
   for now.

No article prose, public checkpoint, social card, PDF, build or upload is
authorized by this audit.

## Required next evidence

Before a full empirical article:

1. obtain the six missing OEIS theorem identities or the authors' counting
   rule;
2. build the problem-level Formal Conjectures baseline at a matching immutable
   revision;
3. design and test a comparable problem-replenishment measure;
4. collect at least one later accepted tracker snapshot without rewriting this
   one;
5. complete broader literature-priority review for the eight retained rows;
6. review the completed limitations and strongest-counterargument documents
   before authorizing any article prose; and
7. keep the Jacobian candidate outside the curve until its independent gate is
   closed.

## Audit verdict

`PFT-2026-07-20` is a valid frozen record of the seed state and reproduces from
its own inputs. The corrected working baseline is eight strict-core Erdős
resolutions from 353 attempted formalized records. The difference is a visible
audit reclassification, not data loss.

The measurement system works precisely because the headline became smaller.
