What Happens When Discovery Starts to Scale?
Earlier this week I posted about OpenAI’s mathematics release. I wanted to know what was in it. Repeating: OpenAI put 722 manuscripts on GitHub on October 6, grouped into 372 families of related results across 17 fields, all from an internal model nobody outside the company can run. By the time of this piece the count was down to 719. OpenAI pulled three a day after release when a sign error took out one argument and two papers that depended on it. [1, 2, 3]
So, I did what I do. I ran deep research passes on the catalog, read a few of the manuscripts I could follow, and bounced the rest off the tools until I understood what was being claimed and what was being checked.
Not everyone is happy about this. A group of mathematicians wrote that nobody asked for the work and called dropping 700 files at once a show of power rather than scholarship. I get it. I’m still going to take the work seriously, because it’s sitting there for anyone to check, and combinatorial mathematician Gil Kalai has already called it a milestone for the field. [3]
I’m not a mathematician. But you don’t need to follow the proofs to see what matters. If a university department published 719 papers in a year, we’d want to know what they were putting in the water. And do some simple math. OpenAI says each result took about three hours of its top-tier thinking compute, on average. Say a serious check of one manuscript takes a working mathematician a week. That’s 719 weeks, roughly 14 years of one person’s time, to read what’s already on the pile. Three hours to make. A week to check. [4]
So, I’m not asking whether the work is real. I’m asking what we’d do differently if this is what research capacity looks like now. Work with me as I try to reframe your thinking.
The factory and the shaft
The best comparison to help everyone picture this in their mind is the factory. We can all imagine one. Now go back in time if you can.
When electric motors showed up, the first thing factory owners did was hang one on the overhead shaft that the steam engine used to turn. Same belts, same layout. Power got a little cheaper. That was the whole benefit.
The real gains came later, when someone gave each machine its own motor and rearranged the floor around how the work moved. That took a long time. Edison’s Pearl Street station opened in 1882. The productivity jump in American manufacturing came in the 1920s. Forty years, give or take, between the technology and the payoff. Warren Devine and Paul David, the Stanford economic historian, both found that the payoff lived in the layout. [6, 7]
Science is laid out the same way
Now look at how we do and have done research.
Every piece of it is a way of rationing one scarce thing: the hours of people who know what they’re doing. Grants decide who gets those hours, and papers are how we package them. A PhD program takes six years to make more. A lab is a building organized around a handful of people whose attention is the limit on everything else.
That is a factory laid out around the shaft. It made sense for as long as expert reasoning was the scarce input. The math release is the first large piece of evidence that it may not stay scarce.
In The Invisible Interface I argued that reasoning is becoming a resource you allocate, the way you allocate capital or compute. (I was writing about companies. I’ll admit I didn’t have mathematicians in mind when I first started writing the book 2 ½ years ago.) This catalog is the first time I’ve seen it play out in research at true scale. If it holds, the question for anyone who runs a research budget changes from how many people you can afford to which questions deserve the reasoning. [8]
What you get if you rearrange the floor
So, what does the rearranged floor look like? Right now, I can see four things from where I sit.
Start with the drawer. Every research group has one, even if nobody calls it that. It’s the set of questions the group set aside because going after them cost too much. Nobody decided those questions were wrong. They were just expensive, and someone had to choose what went forward and what got put to the side. If the cost of an attempt drops far enough, that drawer is the first place to look.
Then there’s the number of attempts. Today a postdoc runs one approach on a problem because that’s all the hours there are. If the AI can run ten in parallel and tell you which three deserve a human’s attention, you learn in months what used to take years.
The bigger change is what happens after a result. Today a mathematician proves a theorem, publishes it, and moves on. That’s the job. Whether the theorem ever turns into something useful depends on a chain of people who may never meet: an engineer who happens to read the paper, someone who builds a method out of it, someone else who tests that method on real measurements and compares it with what people already use. Each of those is a different person with a different grant, often a decade apart. Most of the time the chain breaks somewhere, and nobody notices, because nobody was responsible for the whole thing.
On the new floor, that chain may be one project. Somebody owns it from the proof to the test, the AI does the parts that used to wait for a specialist’s time, and the project ends with a tested method or a written record of why it didn’t work. That’s what I mean when I say discovery stops depending on luck. Today it depends on the right people happening to find each other. It doesn’t have to, if we’re willing to rearrange the floor.
And the failures. Right now, the approaches that didn’t work leave with the person who tried them. They sit on a laptop. The next group starts from zero. If every attempt, including the ones that failed, gets carried into the next attempt, the work compounds. I wrote about that loop in Who Owns the Innovation Loop? This is the loop running inside research. [9]
The answer and the path to it
A proof itself isn’t a yes or a no. It’s a map.
Take one paper that I picked in the catalog. It deals with elasticity, which is how a material deforms when you push on it. The result says that if you can measure everything happening at the surface of an object, then under the conditions the paper lays out, you can work out what it’s like inside. Think about knowing what’s going on inside a bridge beam or a turbine blade without cutting it open. [10]
Read the proof and you see what it leans on. Complete measurements. Ideal conditions. Those assumptions in the proof are the next questions that someone can now ask. What if you can only measure part of the surface, or the readings are noisy? What if there’s already a crack? The answer tells you something is possible. The path tells you where it applies and what you might ask next.
I’ve watched this play out in drug development. A drug that works and nobody knows why is a problem for everyone who comes after it. Without the mechanism you can’t necessarily pick the next indication, you can’t predict who it fails in, and you can’t design the next trial. The result is real. But you want the map.
Now look at what OpenAI released. Abridged reasoning summaries for ten of the 372 families. For everything else you get the result and, for 162 of the manuscripts, a proof in Lean that a computer can check. The advisory group of mathematicians at the Institute for Advanced Study that OpenAI says it consulted had asked labs to disclose the model, the prompts and the compute. OpenAI disclosed the compute. So, for most of the catalog, the answer is there and the path is left out. [4, 5]
We’ve run this experiment before and we should learn from it. In 1976, Kenneth Appel and Wolfgang Haken proved the four-color theorem using 1,200 hours of computer time and roughly 10 billion logical steps that no human could read but turned out to be correct. Regardless of being correct it was resented, because it settled a question that had been open since 1852 without pointing anywhere new. Appel’s own view, as his son told it after he died: “Without computers, we would be stuck only proving theorems that have short proofs.” [11, 12, 13]
Both sides were right. The computer made the proof reachable. A proof nobody can follow doesn’t hand you the next question and that is part of the learning that informs so much.
So, ideally the new floor has to produce paths along with answers. In practice that probably means a second pass, maybe by another AI system, that turns a machine-checked proof into something a person can follow and argue with. It also means the checking never stops, because the output is uneven. Three manuscripts fell off the original list. More will follow. That checking is part of the new floor plan.
What history says you get
Fine. But now let’s ask – what does an organization actually get for reorganizing?
Electricity we covered. The surge came decades after the motor, and it came to the plants that rebuilt the floor.
Computers are the closer case. Erik Brynjolfsson from Stanford and Lorin Hitt followed 527 large US companies from 1987 to 1994 and asked what they got for their computer spending. In the first year, just ordinary returns. Over five to seven years, returns up to five times larger. The difference was the slow, expensive work of reorganizing the company around the machines. That’s the gap between buying the hardware and rebuilding around it. (I cited Brynjolfsson last month on productivity. He’s been on this question for a long time.) [14, 15]
So, the honest answer to what I asked is a multiple, on the condition that you do the reorganizing, and on the understanding that it might take you years depending on how you approach it.
Why this time is bigger (IMHO)
I think the upside this time is larger than either case I cited.
The first reason is that the thing that got cheap is the thing you need to do the redesign. A motor couldn’t help you lay out a factory. Reasoning can help you choose which questions to fund and read the paths behind the answers. The technology is also the tool for adopting it.
The floor is also easier to move. Factories lagged partly because single-story plants had to be built before the new layout was even possible. A research program is made of roles and money. It gets rearranged by decisions about who owns what and what gets funded. Slow, but not forty years slow.
And verified results compound. A motor gave you the same horsepower every day for twenty years. A verified theorem feeds the next attempt, and the failures feed it too.
Now flip it around
Does a company that starts today have to go through any of this? Or does it get a head start because it never had a shaft to begin with?
History says the head start is real. The factories that got the most out of electricity were often the ones built after the motor existed. Nobody had to tear anything out. [7]
And people are already making that bet. Lila Sciences came out of Flagship Pioneering in Cambridge with $200 million in seed money in March 2025, raised a $235 million Series A that September, and signed a 235,500-square-foot lab lease, one of the largest in Greater Boston, to build what it calls AI science factories. Periodic Labs, started by Liam Fedus, who ran post-training at OpenAI, and Ekin Dogus Cubuk from Google DeepMind, raised a $300 million seed round to build autonomous labs for materials. Its own launch note points out that those labs generate negative results that seldom get published. That’s the failures paragraph above, as a business plan. [17, 18, 19]
Axiom Math is the closest to this release. Its system writes proofs in Lean, the language computers can check, and then its human mathematicians write the explanation to go with each one. Ken Ono, the number theorist who joined as founding mathematician, told Axios that five journals have accepted papers built that way. Formal proof plus a path a person can follow. That’s the second pass I described. [20, 21]
So yes, a new organization skips the rearrangement. That said – It still needs judgment, and judgment is the part nobody can build from scratch. The people who can look at 719 results and tell you which five matter were trained inside the old system, and every one of these companies is planning on hiring them. The head start is real, and it’s bought with people from the incumbents.
It’s bought with compute too. I wrote in Compute about young AI companies that can’t get the capacity they’re willing to pay for. Run the number here. Three hours of top-tier compute per result, times 372 results, is more than 1,100 hours, about 46 days of a frontier model running flat out, before anyone checks anything. A university doesn’t have that line item. A startup must finance it before it has a result to show. [16]
The incumbents have something the new entrants don’t: the drawer of unattempted. Thirty years of questions somebody set aside, plus the data, the instruments and the benches. A startup’s drawer may be empty on day one. My guess is new entrants win first where the test is cheap, math and software, and incumbents keep the advantage where the test needs a wet lab or a reactor, unless they sit on it. As an investor, the first thing I’d ask a new entrant is where its ideas come from.
Where it shows up first
This will land first wherever a result can be checked by running it. Software is the obvious case. Chip design is close behind, because a design can be simulated before anything gets fabricated. (building and scaling a chip is a different article) Physics models come next, then chemistry, because a lab can test what the math predicts. Biology always comes last, and I say that after thirty years. The testing is slow and expensive, and what you learned in a mouse often doesn’t carry into a person.
Most organizations may hang this new capability on the shaft they have internally. They’ll run the model against the same grant cycle and the same publication count, get a modest gain, and conclude that’s what it was worth or not. A few who can break the existing frame are going to rearrange the floor and find out what it’s worth.
Two questions for your next review. Which questions did we set aside because of cost, and would we make the same call today? And who owns a result after it’s published?
If anyone is already running research this way, I would love to talk to you about it.
I have so many other thoughts in my head about this shift – just let your imagination go and many pieces come into play if what I say above is the new paradigm.
Harry Glorikian is the author of The Invisible Interface: How AI Turns Intentions Into Actions—And Who Wins (Ideapress Publishing / Simon & Schuster, June 2026). He is General Partner at Scientia Ventures, an Affiliate Researcher at the MIT Media Lab, and host of The Harry Glorikian Show.
References
[1] Harry Glorikian, LinkedIn post on the OpenAI mathematics release, October 7, 2026. https://www.linkedin.com/feed/update/urn:li:activity:7513614489468985344/
[2] OpenAI, math repository, released October 6, 2026. https://github.com/openai/math
[3] Davide Castelvecchi, "OpenAI posts 700 maths preprints online: mathematicians are up in arms," Nature, October 7, 2026. https://www.nature.com/articles/d41586-026-03196-8; Alicia Gallegos, "OpenAI withdraws three preprints a day after releasing 722 manuscripts on unsolved math problems," Retraction Watch, October 8, 2026 (the three withdrawals and the Association for Human Mathematics statement). https://retractionwatch.com/2026/10/08/openai-withdraws-preprints-722-manuscripts-unsolved-math-problems/; Gil Kalai comment via 36Kr. https://eu.36kr.com/en/p/4017976429907840
[4] Decrypt, October 2026: 4,000 problems posed, three hours of compute per result, 162 Lean formalizations, ten reasoning summaries, OpenAI’s own caveat. https://decrypt.co/380366/openai-secret-ai-model-cracked-hundreds-math-problems-one-prompt
[5] IAS Advisory Group on Mathematics and AI recommendations, September 29, 2026, as summarized by Emergent Mind; OpenAI’s omission of model name and prompts per Let’s Data Science. https://www.emergentmind.com/openai-math-explorer and https://letsdatascience.com/blog/openai-posted-722-ai-written-math-papers-kept-the-prompts
[6] Warren D. Devine Jr., “From Shafts to Wires: Historical Perspective on Electrification,” Journal of Economic History 43(2), 1983.
[7] Paul A. David, “The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox,” American Economic Review 80(2), 1990.
[8] Harry Glorikian, The Invisible Interface, Ideapress Publishing, 2026. https://www.amazon.com/dp/1646872487 and https://glorikian.com/invisible-interface/
[9] Harry Glorikian, “Who Owns the Innovation Loop?” August 4, 2026. https://glorikian.com/who-owns-the-innovation-loop/
[10] Elasticity manuscript, OpenAI math repository.
[11] Boston Globe, “Kenneth Appel, 80; first to use a computer to prove a major math theorem,” April 29, 2013. https://www.bostonglobe.com/metro/obituaries/2013/04/29/kenneth-appel-dies-used-computer-map-question/nEKrYrHng1gTWI0iUxVMON/story.html
[12] UPI, “Mathematician Kenneth Appel dies at 80,” April 30, 2013. https://www.upi.com/blog/2013/04/30/Mathematician-Kenneth-Appel-dies-at-80/2261367349696/
[13] Nashua Telegraph, “Proving the four-color theorem with computers,” April 29, 2013. https://www.nashuatelegraph.com/news/local-news/2013/04/29/proving-the-four-color-theorem-with-computers-made-the-most-famous-mathematician-in-nh-history
[14] Erik Brynjolfsson and Lorin M. Hitt, “Computing Productivity: Firm-Level Evidence,” Review of Economics and Statistics 85(4), 2003. https://doi.org/10.1162/003465303772815736
[15] Harry Glorikian, “The AI Slowdown Debate Has a Surrogate Endpoint Problem,” September 16, 2026. https://glorikian.com/the-ai-slowdown-debate-has-a-surrogate-endpoint-problem/
[16] Harry Glorikian, “Compute,” September 10, 2026. https://glorikian.com/compute/
[17] Reuters via Yahoo Finance, “AI lab Lila Sciences tops $1.3 billion valuation with new Nvidia backing.” https://finance.yahoo.com/news/exclusive-ai-lab-lila-sciences-101645837.html
[18] Fierce Biotech, “Flagship’s Lila Sciences lands $235M,” September 2025. https://www.fiercebiotech.com/biotech/flagships-lila-sciences-lands-235m-expand-ai-powered-autonomous-research-labs
[19] Periodic Labs launch note, via Marginal Revolution, October 2025. https://marginalrevolution.com/marginalrevolution/2025/10/ai-scientists-in-the-lab.html; Maginative, “Periodic Labs launches with $300M.” https://www.maginative.com/article/periodic-labs-launches-with-300m-to-build-an-ai-scientist/
[20] Axios, “AI math startup’s proofs land in peer-reviewed journals,” May 26, 2026. https://axios.com/2026/05/26/axiom-ai-math-journal
[21] Seedtable, Axiom Math Series A, March 2026. https://seedtable.com/companies/axiom-math/funding-rounds/series-a-2026-03
If this resonated with you
The Invisible Interface
My new book explores how AI is creating an invisible operating layer between businesses and their customers — and what that means for every leader making strategy decisions today. Available now.

