A year of work, a rumor, and 88 hours

Two mathematicians ran a year of unpublished reasoning through someone else's API. When a competing result appeared eight days after a rumor, the only honest answer anyone had was a shrug.

September 9, 2026 9 min read
Why this matters to non-mathematicians Most work doesn't need a frontier model What makes self-hosting affordable Where self-hosting is the wrong answer What you could actually prove

On September 8, 2026, OpenAI published a proposed resolution to the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize problems. Nobody had solved it in about ninety years. An unreleased internal model running roughly 10,000 concurrent agents in communicating groups got there in 88 hours.

That got quoted everywhere. What happened around it got much less attention, and it should matter more to anyone running an engineering organization.

A sunlit university blackboard densely covered with handwritten mathematical equations, some sections partly erased

Tristan Buckmaster at NYU and Levent Alpöge, a mathematician at Anthropic, had spent close to a year working a related route into the same territory. They used Codex and Claude heavily the whole way. On August 15 they proved finite-time blow-up for the Euler equations, the frictionless cousin of Navier-Stokes, and verified it in Lean. While they were writing it up, they learned that word of their approach had reached OpenAI.

OpenAI says its own effort began September 1, prompted by a rumor that two mathematicians were closing in. Its researchers state that neither they nor the agents saw any of Buckmaster and Alpöge's work before it went public, and that no specific user data was accessed to solve the problem. Their statement also carries a line that is easy to skim past: the company says it "cannot rule out" that de-identified data from the pair's use of its products helped improve its models generally.

~ 1 year

Two mathematicians work a related route

Buckmaster (NYU) and Alpöge (Anthropic) push toward the same territory, using Codex and Claude heavily the whole way.

Aug 15

Finite-time blow-up for Euler, verified in Lean

They prove the result for the frictionless cousin of Navier-Stokes and check it in Lean. While writing it up, they learn word of their approach reached OpenAI.

Sep 1

OpenAI begins, on a rumor

By its own account, OpenAI starts after hearing that two mathematicians were closing in. It says the agents saw none of the pair's work before it went public.

Sep 8 · 88 hrs

OpenAI publishes a proposed resolution

An internal model running ~10,000 concurrent agents reaches a proposed Navier-Stokes resolution in 88 hours - eight days after the rumor started moving.

Buckmaster has been careful here. He has said he is "not accusing anyone of anything" on that specific question. He does not know whether their private sessions were used, and neither does anyone outside the company.

That unresolvable quality is the part I keep coming back to. Two mathematicians ran a year of unpublished, career-defining reasoning through someone else's API. When a competing result appeared eight days after a rumor started moving, the only honest answer available to anybody was a shrug. The provider could not rule it out, the researchers could not rule it in, and there is no way to go back and check. Once those prompts left their machines, no record existed that either side could put in front of a referee.

“

Once those prompts left their machines, no record existed that either side could put in front of a referee.

Why this should bother people who are not mathematicians

Swap the nouns and the story is about your company. Your pricing model instead of a proof. Your clinical trial protocol, your acquisition shortlist, the claims language your legal team spent nine months negotiating. Route any of that through an external endpoint and you inherit the same shrug.

The counterargument is familiar and mostly fair. Enterprise agreements exclude your data from training, zero retention is available, and for plenty of workloads that really is enough.

But look at what this fight actually turned on. There was no breach and no violated contract. A rumor reached a competitor, a team got pointed at a problem, and the ambiguity did the rest. A contract protects you in court, and it does very little when the question is how somebody knew, because you have no records of your own to show anyone and you end up arguing about it in public.

“

A contract protects you in court. It does very little when the question is how somebody knew.

Two other things landed the same week that are worth reading next to this one. OpenAI shipped ChatGPT Images 2.5 on September 8, cutting generation latency by up to half on a product already producing more than three billion images a week. Separately, Sony AI's scientific discovery team has been publishing on literature-based hypothesis generation, using temporal knowledge graphs to surface gene and disease links that the published record implies but has not yet stated.

The detail worth borrowing

Sony AI did not have a frontier lab's compute budget, and their senior research scientist Uchenna Akujuobi credits that constraint with pushing the team toward approaches that were computationally lean and still scientifically effective. Their models also let a researcher walk back through the graph and see where any prediction came from - which you cannot do with a frontier API call. The compute limits made the engineering better.

The compute limits made the engineering better, which is worth remembering the next time somebody tells you the biggest available model is the safe default.

Most of your work does not need a frontier model

OpenAI's run reportedly consumed around 300 billion output tokens. At current rates that puts it somewhere near $22.5 million of compute for a single proof.

~10,000
concurrent agents in communicating groups
~300B
output tokens reportedly consumed
~$22.5M
compute for a single proof, at current rates

I don't read that as a story about frontier labs. It says more about how most companies build. Teams reach for the largest available model by default, route every request to it, and then act surprised at the invoice. Classification, extraction, routing, summarization, tagging, redaction, form parsing: that is the real shape of enterprise AI work, and almost none of it needs a frontier model.

“

A fine-tuned 7B model that has seen ten thousand of your own examples will usually beat a general-purpose giant on your narrow task, run on hardware you can name, and cost about as much to train as lunch.

On the Strongly.AI platform, a 7B LoRA job over 1,000 examples for 3 epochs on an NVIDIA T4 runs about an hour and lands near a dollar. A 13B QLoRA run over 5,000 examples comes in around five. A 70B QLoRA job over 10,000 examples is roughly twenty.

Fine-tune job
Examples
Notes
~ Cost
7B LoRA
1,000 × 3 ep
~1 hr on T4
~$1
13B QLoRA
5,000
single job
~$5
70B QLoRA
10,000
single job
~$20
Frontier proof run
~300B tokens
one problem
~$22.5M

A dollar against twenty-two million is not a fair comparison - they are obviously different problems. But most of what your business actually runs on looks a lot closer to the first row.

A dollar against twenty-two million is not a fair comparison, since they are obviously different problems. But most of what your business actually runs on looks a lot closer to the first one.

What makes self-hosting affordable

Owning the weights is the easy part. Self-hosting tends to fail inside organizations for a duller reason: a GPU left running is a GPU you pay for at three in the morning while nobody is awake to use it. Sovereignty that costs four times as much does not survive its first budget review.

The Strongly.AI AI Gateway is built around that specific failure. Self-hosted models deploy from Hugging Face onto your own infrastructure, served by vLLM, Transformers, or a custom image, inside your organization's own Kubernetes namespace. After that the economics are yours to set.

Always On

Keeps a model warm for production traffic with no cold starts. The right default only for the workloads that actually run all day.

70-90% savings On Demand

Scales to zero after a configurable idle window and wakes on the next request. The platform's own guidance puts this at 70 to 90 percent savings for intermittent workloads.

~1/4 the cost Scheduled

Runs on a cron expression. A model that only serves the London office can live on 0 9-17 * * 1-5 and cost roughly a quarter of a round-the-clock deployment.

up to 70% cheaper Autoscaling + Spot

Sizes replicas against live in-flight request counts, scaling up immediately and down with hysteresis so pods don't flap. Spot capacity sits on top, up to 70% cheaper, with automatic fallback to on-demand.

Start with deployment mode. Always On keeps a model warm for production traffic with no cold starts. On Demand scales to zero after a configurable idle window and wakes on the next request, which the platform's own guidance puts at 70 to 90 percent savings for intermittent workloads. Scheduled runs on a cron expression, so a model that only serves the London office can live on 0 9-17 * * 1-5 and cost roughly a quarter of a round-the-clock deployment.

Autoscaling covers the traffic you cannot predict. On-demand deployments can combine scale-to-zero with demand-driven replica scaling, sizing replicas against live in-flight request counts rather than lagging CPU metrics, scaling up immediately and scaling down with hysteresis so GPU pods do not flap. Spot capacity sits on top of that, up to 70 percent cheaper, with automatic fallback to on-demand so you are not trading availability for price.

Routing gets used less than it should. The gateway fronts 316 certified models across 12 providers behind one OpenAI-compatible endpoint, alongside anything you host yourself, so task complexity rather than habit decides where a request goes. Simple classification hits your fine-tuned 7B. Ambiguous edge cases escalate to a frontier model. Anything touching material you cannot afford to leak never leaves your namespace. Strongly's documentation puts routing savings at 50 to 80 percent, though the number I would watch is the share of sensitive traffic that stops crossing your perimeter.

316
certified models behind one OpenAI-compatible endpoint
12
providers, alongside anything you host yourself
50-80%
documented routing savings - and less sensitive traffic leaving your perimeter

Guardrails run on the completions path, covering content filtering, PII detection, and prompt injection protection. Deployments are scoped to organization namespaces with tenant isolation, and every call carries token counts, latency, and cost attribution you can query by model, provider, or user.

“

Buckmaster and Alpöge had none of that. What would have helped them was their own logs, not better contract language.

Where self-hosting is the wrong answer

Said plainly

Strongly's own cost documentation is blunt about this. For most usage patterns, third-party providers stay cheaper until you reach very high volume. Below roughly a billion tokens a month a hosted API usually wins on price, and self-hosting only becomes competitive with real optimization and sustained load.

Strongly's own cost documentation is blunt about this, and it deserves repeating rather than burying. For most usage patterns, third-party providers stay cheaper until you reach very high volume. Below roughly a billion tokens a month a hosted API usually wins on price, and self-hosting only becomes competitive with real optimization and sustained load.

The actual argument

So the argument is not really about cost. Cost is what makes control affordable enough to keep.

The right architecture for almost everyone is mixed. Frontier APIs for genuinely hard reasoning on material you don't mind sending out. Fine-tuned small models on your own infrastructure for the high-volume work and for anything you would not want to explain in a deposition. One gateway in front of both, so routing policy stays a configuration change instead of a rewrite.

What you could actually prove

The mathematicians will be fine. Their proof is verified in Lean, and Lean does not care who published first.

Your position is weaker than theirs. You will not get a Millennium Prize problem and a public statement from a frontier lab. You will get a competitor who moves earlier than they should have been able to, plus a perfectly plausible explanation in which nobody did anything improper.

“

If the work you are running through an external model turned up somewhere else in six months, what could you actually prove?

So the question for your next architecture review is not whether your provider is trustworthy. It is narrower than that. If the work you are running through an external model turned up somewhere else in six months, what could you actually prove?

If the answer is nothing, that is worth fixing while it is still an architecture decision rather than an incident review.

Keep the work where you can prove it

The Strongly.AI AI Gateway supports self-hosted deployment from Hugging Face with autoscaling, scheduled and on-demand modes, fine-tuning with LoRA, QLoRA, full and prompt tuning across 30+ open-source base models, guardrails, and unified routing across 316 certified models and 12 providers.

Explore the AI Gateway