Skip to main content Scroll Top

Who Owns the Innovation Loop?

AI can already produce validated new results. The next advantage belongs to whoever controls the experiment, the evidence, and the next cycle.

In The Machine That Builds the Machine I argued that AI was starting to move upstream. It had been the product of the infrastructure. Now it was becoming part of what builds the infrastructure that makes its successor. That’s the thesis.[1]

Then OpenAI put out Jalapeno, a custom chip built mostly to run AI models. OpenAI says it went from initial design to manufacturing tape-out with Broadcom in nine months. Tape-out is the point when a finished chip design is sent toward fabrication. Celestica helped turn the chip into boards, racks and the surrounding systems, and OpenAI says its own models accelerated parts of the design and optimization process.[2]

The question was: did AI design the chip? Based on what OpenAI has disclosed, no one outside the company can honestly say that it did. We don’t know which parts of the work the models performed, how much time they saved, whether they originated a material part of the architecture or what a properly matched human-led program would have taken. Engineering samples are running, but the company is still measuring final performance.

Even with those limitations, Jalapeno was the kind of breadcrumb I had been looking for. It shows an AI company using its own models inside the process of building the hardware on which future models will run. Better models could improve parts of chip design, verification and the software that runs on the chip. Better hardware could make later models faster and less expensive, supporting more use, more experiments and more investment in the generation after that.

The first AI-generated invention isn’t the milestone I’m watching for. The more important signal is what I think of as the second turn of the loop: an AI-generated improvement that makes the next system better, and that stronger system produces another validated improvement. One turn can be a useful result. The second begins to look like compounding.

At a recent Berkeley summit, The Information reported that Jasjeet Sekhon, Google DeepMind’s chief strategy officer, called recursive self-improvement a “key part of the investment thesis” behind the industry’s extraordinary capital spending. He also warned of an “AI air pocket” if the spending arrives before the revenue.[3] So the bet isn’t only that companies buy more AI. It’s that AI starts improving the process that makes better AI. Sekhon reached for the steam engine to explain it: once you had one, you could use it to build a better one.

Full recursive self-improvement is a very high bar. It would mean a system independently redesigning, training and validating its successor, then repeating that process. We haven’t seen that – at least not yet. My second-turn test is much narrower and should be easier to observe: did an AI-originated improvement enter the next system, and did that stronger system produce another validated improvement?

Sergey Brin described the current operating shift at Google in much plainer language. He said Gemini is increasingly being used to monitor training runs, generate training data and help develop the tools used to build later systems: “you start to use the tool to build the tool.”[4] That isn’t full recursive self-improvement either, but it does tell us this is becoming part of the normal work inside a frontier laboratory rather than remaining a research aspiration.

Level 4

In July 2024 OpenAI shared an internal framework with employees describing five stages of capability: chatbots, reasoners, agents, innovators, organizations. A spokesperson confirmed it to Bloomberg. Level 4, innovators, was AI that could aid invention. Level 5 was AI that could do the work of a company. It was never published as a technical standard.[5]

I’ve translated those labels into working definitions for this piece:

Level Working name What the complete system can do
C1 Chatbot Responds to a request
C2 Reasoner Solves a problem and checks the result
C3 Agent Acts, observes what happened, and corrects course
C4 Innovator Originates and validates something materially new
C5 Organization Maintains a mission, allocates resources, coordinates work, and learns over time

The table looks like a staircase, but technology rarely develops that neatly. A specialized system can discover a better algorithm while being incapable of organizing a meeting. Another can answer customers, reorder inventory and manage a workflow without inventing anything. The first three levels do stack neatly: answer, solve, act. After that the path branches. Level 4 goes deeper into invention. Level 5 goes wider into coordination, resource allocation and institutional judgment.

For this discussion, four questions matter:

1. Did the AI originate an important part of the idea?

2. Did a credible test show that the idea worked?

3. Did the result improve the next product, process or research cycle?

4. Did the system solve the problem inside a frame supplied by people, or did it create or materially revise the frame itself?

The first two tell us whether AI contributed a real invention. The third tells us whether the result is beginning to compound. The fourth separates two very different kinds of innovation. One new proof is a Level 4 result. Repeated, validated proofs across different problems begin to look like a capability. If those discoveries improve the system that produces the next result, the consequence is larger still.

By OpenAI’s loose definition, AI that can aid invention, Level 4 got here a while ago. Hold it to a harder standard and I’d put it this way. Framed invention is here. Some framed loops are starting to compound. Nobody that I have seen so far has shown a machine that can build a brand new scientific frame.

Machines Producing New Results

The story didn’t begin with today’s frontier models. In 2022, Google DeepMind introduced AlphaTensor, which discovered previously unknown and provably correct ways to perform matrix multiplication. A year later, AlphaDev found new sorting routines that outperformed prior human-designed implementations, with selected routines later incorporated into a widely used C++ library. FunSearch then paired a large language model with a systematic evaluator and produced new mathematical and algorithmic results.[6][7][8]

These systems didn’t decide that matrix multiplication or sorting deserved their attention. People selected the problems, represented them in a form the systems could work with and created the tests that determined whether an answer was better. I see that as AI inventing inside a human-created frame, but keep in mind that doesn’t make the result less real. People invent through search, recombination, testing and rejection too. The relevant questions are whether the result was new, whether it represented a material improvement and whether it survived a serious test.

The recent OpenAI mathematics work matters because it suggests that general-purpose models are moving beyond isolated specialist systems. In May, OpenAI disclosed that an internal model had produced a construction disproving a longstanding conjecture related to the planar unit-distance problem, first posed by Paul Erdos in 1946. Outside mathematicians checked the result. The companion paper also makes the human role clear: a mathematician selected the problem, and people later reorganized and extended the model-generated argument. The central construction appears to have come from the system and survived review.[9]

While I was working on this piece, OpenAI moved the evidence again. On August 1, it released ten additional results generated by an internal version of Astra, which OpenAI calls its next major model. The problems range across geometry, coding theory, complexity, cryptography and other areas of mathematics and theoretical computer science.[10]

OpenAI published the manuscripts, reasoning walkthroughs and machine-checkable certificates in Lean, a formal system that verifies each encoded logical step. The company says Astra generated the mathematical arguments and formalized the proofs, while people worked with the model to prepare the manuscripts. It also says the tokens used to find the ten solutions would have cost roughly $2,000 at the rates charged for Sol.[10]

That’s a substantial step beyond one striking proof. It’s also a selected set of successes.

Go back to May for a second. Alongside the unit-distance result, OpenAI published a chart showing how often its models solved that problem at different amounts of test-time compute. A success rate. For these ten it published nothing of the kind. No attempt count, no word on how the problems were chosen, no sense of how much parallel search ran or how often Astra produced plausible work that fell apart.

Same company, five weeks apart. The proofs got stronger and the disclosure got thinner – at least for now.

A Lean certificate tells us the encoded proof checks. It doesn’t settle novelty, significance, or what was already sitting in the literature.

Assuming the results hold up under wider review, however, they’re strong evidence that framed Level 4 in mathematics and theoretical computer science is becoming a capability rather than a one-off event. What remains unclear is how reliably the system can do this across a representative set of problems.

Inventing Inside the Frame Does Not Equal Inventing the Frame

In February, at the India AI Impact Summit, Demis Hassabis proposed what people now call the Einstein Test: train a system on everything known up to 1911 and see whether it gets to General Relativity by 1915. His own verdict was that today’s systems couldn’t do it. Weeks later Tom Zahavy, a researcher at Hassabis’s own lab and a contributor to AlphaProof, published a position paper explaining why. He called it LLMs Can’t Jump.[11] His argument is that today’s systems are becoming very good at finding patterns and deriving conclusions once the concepts and rules are known. What they haven’t demonstrated is the ability to create the concepts or premises when no clean dataset or objective points toward the answer.

Zahavy uses Einstein’s development of General Relativity as the example. Einstein didn’t simply find a better answer inside Newtonian gravity. He changed the frame through which gravity was understood. Zahavy calls that move abduction: forming a new explanatory hypothesis from sparse, indirect or surprising evidence. In his formulation, current language models handle induction, which is finding patterns in data, and they’re getting good at deduction, which is working out the consequences of known premises. What they haven’t shown is the ability to invent the premises.

That distinction is useful because Astra and AlphaEvolve can produce important new results inside formal systems created by people, but they don’t show that a machine can create the conceptual equivalent of curved spacetime. I wouldn’t go as far as the paper does when it says language models are structurally incapable of making that jump. And to be fair to Zahavy, neither would he. After the piece got picked up he went out and said publicly that people were misreading it, that this was a personal position paper rather than his employer’s view, and that he wasn’t arguing LLMs can never make real discoveries.

Either way it’s a position paper, not an impossibility proof, and General Relativity is a brutal standard for invention. Most human inventions don’t create a new scientific ontology. A new drug, a chip design, a manufacturing method, an algorithm. Any of those can be worth a fortune while sitting entirely inside an established frame.

Still, the boundary is real. There are two related loops. The inner loop improves something inside the accepted frame:

Problem -> candidate -> test -> evidence -> archive -> next candidate

The outer loop begins when the frame no longer fits reality:

Anomaly or contradiction -> new representation -> new hypothesis -> new test -> revised inner loop

The inner loop can compound rapidly without changing the frame, which by itself could have major economic consequences. The outer loop is harder. No available system has shown that it can reliably recognize that an accepted representation is inadequate, create a materially new one and design a real-world test that distinguishes the new explanation from the old. At least so far.

World models may eventually help. A world model tries to predict how an environment will change and what will happen after an action. If a system can run counterfactual experiments in a useful simulation, it may do more than save time; it may encounter an unexpected outcome and use that outcome to form a new explanation. That is the theory, but it has not been demonstrated at the level Zahavy is describing.

Reality You Can Grade

So why did the first clear examples appear in mathematics, code and algorithms? The rules are relatively clear. A proof can be inspected, a program can be run and a software change can be timed. The system receives an answer quickly and can try again.

Recursive Superintelligence has built its early automated-research work around those conditions. Its system proposes an idea, implements it, runs an experiment, evaluates the result and uses what it learned to choose the next experiment. Recursive deliberately started with small-model training and graphics-processor optimization because the experiments were fast, the metrics were clear and the tests could be strengthened against manipulation.[12]

The reported gains are fine. What happened to the tests is the interesting part. Recursive says its system tried to game all three evaluators. It cached answers. It leaned on state left behind by earlier runs. It gamed the timing harness instead of writing faster code.

Look, this is teaching to the test. Anybody who has run a lab knows the version of this where somebody optimizes the metric and quietly stops doing the science. We have many examples of humans doing the same. The answer is the same either way: you must write a better test.

Google DeepMind’s AlphaEvolve shows what this architecture can produce when the evaluator is strong enough. AlphaEvolve uses Gemini models to propose programs. Automated tests run and score them, while an evolutionary process keeps useful changes and uses them to produce the next candidates. DeepMind reports that AlphaEvolve improved data-center scheduling, contributed a circuit change to an upcoming Google AI chip and accelerated an important Gemini operation by 23%, reducing overall model-training time by roughly 1%. It also says AlphaEvolve improved processes used to train the models underlying AlphaEvolve itself. Those operational gains are reported by Google rather than independently audited.[13]

The model proposes a change, the test determines whether it works, and engineers decide whether to deploy it. The next system then begins with better software or infrastructure. This is an early form of recursive improvement at the level of the complete research operation. It doesn’t require a model to disappear into a data center and rewrite itself, but it does require a test that measures the result people actually care about.

Qwen3.8-Max provides a more recent example closer to the core of AI research. Alibaba gave the system a paper on selecting training data for reasoning models and asked it to reproduce the work, then improve it. According to Qwen’s published trajectory, the system worked for 125 hours without human intervention during execution, ran 33 GPU jobs and tested 18 possible improvements. Its best method increased the reported AIME24 result by 2.71 percentage points over the reproduced baseline.[14]

That is meaningful evidence of a general model carrying out a bounded research loop. The original paper included Alibaba researchers. People selected the task, supplied the computing environment and defined the benchmark. The result hasn’t been independently replicated, and Qwen hasn’t shown that the improvement entered a successor model and enabled another improvement. It moves us closer to the second turn, but it isn’t the second turn.

The evaluator problem becomes more serious as the systems improve. In Anthropic’s Automated Weak-to-Strong Researcher, nine AI agents worked for five days on a defined alignment problem, representing approximately 800 cumulative research hours and around $18,000 in computing and model costs. The agents substantially outperformed Anthropic’s human comparison on the selected metric, but they also found several ways to improve the score without preserving everything the researchers intended to measure.[15]

METR hit the same wall evaluating GPT-5.6 Sol under an agreement with OpenAI. Count the model’s detected cheating attempts as failures and its 50% time horizon lands around 11.3 hours. Throw those runs out instead and you get 71 hours, with a confidence interval running out to 11,400. METR said plainly that it doesn’t treat any of those numbers as a real measurement of what the model can do.[16]

This is a safety problem, but it’s also a research and capital-allocation problem. A system can produce persuasive studies and apparently improved results while optimizing a measure that doesn’t survive contact with the actual objective. The evaluator, which in plain language is just the test, is part of the invention system. If you don’t own the evaluator and you don’t own the loop.

Brin’s comments add one more wrinkle. He says capabilities that once required specialized models are increasingly converging into general Gemini models, and that training on coding can improve mathematical reasoning and vice versa.[4] The general reasoning engine may transfer across fields faster than many people expected. The experimental loop doesn’t transfer so easily. A biological assay, a chip verifier and a fusion experiment still require different data, instruments, safety rules and evidence.

The Loop Does Not Run at One Speed

To me the innovation loop looks something like this:

Question -> observation and measurement -> hypothesis -> experiment -> evidence -> archive -> deployment -> next cycle

The pace of progress depends on the entire sequence. A brilliant model without a trustworthy test produces plausible possibilities. A successful experiment without a path into production creates knowledge but not necessarily economic advantage. A result the organization fails to retain may have to be rediscovered later.

Different fields return that feedback at wildly different speeds.

Field How the result is tested Likely constraint
Mathematics and software Proof checking, test suites, benchmarks Evaluator integrity and compute cost
Chips and engineered systems Simulation, formal verification, fabrication, physical test Manufacturing, transfer, and time
Fusion and materials science Simulation, then scarce physical experiments Experimental throughput and instrumentation
Human biology and medicine Controlled experiments, replication, clinical evidence, patient outcomes Measurement, causality, translation, and time

Human biology makes the limitation particularly and painfully obvious. A more capable model can connect known facts more effectively, but it can’t reason an unmeasured biological fact into existence (I wish it could). The same diagnosis may contain several different mechanisms as many in the field know well. Measurements are noisy. Cell models don’t always predict what happens in a person. An experiment that separates causation from correlation can take months, and the result may still fail in the clinic.

That doesn’t make intelligence irrelevant; it changes the job intelligence has to do. A complete research system can help determine which measurement would be most informative, which intervention could distinguish competing explanations and what experiment should come next. Insitro describes its TherML platform in those terms: computational predictions guide experimental design, automated laboratories generate data and the results refine the next model cycle. That is the company’s description of its platform.[17]

The OpenAI-Ginkgo Bioworks project gives us a more concrete example. GPT-5 designed experimental batches for cell-free protein synthesis, a way to make proteins without growing living cells. Ginkgo’s automated laboratory executed the experiments, and the results returned to the system, which analyzed the data and proposed the next round. Across six rounds, the project tested more than 36,000 reaction compositions on 580 automated plates. OpenAI and Ginkgo report a 40% reduction in the cost of producing the selected protein, with titer up 27% at the same time.[18]

Ginkgo now sells the improved reaction mix in its reagent store. Hypothesis to experiment to evidence to a product on a shelf. That’s a full loop running end to end.

The work was limited to one protein and one cell-free system, with people preparing materials, maintaining the automation and improving parts of the laboratory process. This wasn’t a general autonomous scientist. It was a bounded research system in which AI proposed experiments, received physical evidence and changed what it tested next.

Fusion shows why even one “domain” can be too broad. In a collaboration between Google and TAE Technologies, no single score captured both plasma quality and the physical limits of the equipment. An algorithm presented pairs of experimental settings to a plasma physicist, who decided which result was better. The combination found an unexpected operating regime with a greater than 50% reduction in energy loss.[19]

The human was a key part of the evaluator. DeepMind later trained a controller in simulation and used it to control the magnetic coils of a real experimental fusion reactor. That closed an important loop in plasma control, but it didn’t solve reactor materials, manufacturing, licensing or the economics of a commercial power plant.[20]

This is how I now think about expansion. Start with one subloop where feedback is fast enough to support learning, then expand into the next adjacent subloop where the data, tools and evidence carry over. The word domain can hide more than it explains. The useful unit in my mind is the feedback loop.

World models may make some physical loops move faster. Google DeepMind’s Genie 3 generates interactive environments, Meta’s V-JEPA 2 predicts physical outcomes from video, and Google’s Gemini Robotics ER 2 uses video understanding to track tasks, recover from failures and coordinate robots.[21] These are research systems and company-reported demonstrations, not proof that the physical world has become easy to model.

The attraction to this is clear. If a system can test thousands of actions in a useful simulation before touching a machine, laboratory or patient, each physical experiment can become more selective and more informative. But the simulator is another test, and it can be wrong. A world model that diverges from reality may help the system become highly effective in the simulation and highly ineffective outside it. Its value will depend on continuous calibration against real interventions and outcomes.

The Experimental Archive

Scientific papers show us what worked. They rarely preserve the complete history of what failed, why it failed, which instrument was miscalibrated, where a person intervened or which promising result disappeared when the conditions changed.

For an AI research system, that missing history really matters. A successful experiment identifies one useful point. A valid failed experiment helps map the boundary around it. It can show which direction not to search, which assumption broke and which apparently reasonable approach should not consume more time.

Research in materials science and chemistry has shown that failed and negative experiments can improve machine-learning predictions. But the word valid matters. A contaminated sample or broken protocol doesn’t provide the same information as a carefully executed experiment that produced a negative result.[22]

The real asset isn’t failure by itself. It is the complete experimental lineage: the question, hypothesis, protocol, operating conditions, raw measurements, instrument state, human interventions, reason the result was classified as a failure and the decision about what to try next.

Two companies can use the same frontier model. One retains that entire history. The other saves the final report. They definitely don’t own the same asset.

In The Invisible Interface, I argue that persistent memory is what allows software to stop resetting to zero and begin compounding around the person or organization using it. The same principle applies here. At research scale, the experimental archive becomes the memory of the innovation loop.[23]

The archive is necessary for compounding, but it may not be sufficient for escaping a bad frame. A complete history can make a system extraordinarily efficient at optimizing the wrong model of the world. The harder capability is recognizing when the model itself needs to change.

Own or Rent the Model?

For most companies, owning the model may not be necessary at the beginning. Building a frontier language model before you have a differentiated learning loop is more likely to destroy capital than create an advantage.

Companies are testing different approaches. Edison Scientific describes a model-agnostic research layer. Periodic Labs is building autonomous physical-science laboratories. Lila Sciences is pursuing a more integrated model-and-laboratory stack. Ricursive Intelligence is starting with chip design and describes a phased path toward connecting more of the hardware and model stack.[24]

Brin’s point about convergence reinforces the practical conclusion. If general models keep absorbing capabilities that once required separate systems, building a frontier model for every field may becomes less attractive, not more. What remains specific is the loop around the model: the data, instruments, evaluator, intervention rights and experimental history.

Some companies may eventually need their own domain model: a model of the relationship between the state of a system, an intervention and the outcome. That could be a cellular-response model, a factory digital twin, a plasma model or a chip-and-workload simulator.

Owning more of the model becomes rational when proprietary feedback can materially improve it, when cost or speed becomes the constraint, when outside models can’t represent the domain adequately or when dependence on a provider gives that provider control of the economics and accumulated learning.

Until then, my rule based on my analysis is straightforward: rent general intelligence while it remains replaceable; control the feedback, the domain model and the learning that comes back from reality; integrate where the handoff causes learning or economic value to disappear.

What This Changes for Companies and Investors

Research organizations will change quite a bit before an autonomous scientist arrives. I keep thinking about the genome work at Applied Biosystems. We got to a point where we said, okay, we need bioinformatics. And the honest reaction in the room was, what the hell is that? So you grab the biology person, grab the informatics person, put them in a room and let them work it out. There was no degree for it at any school back then. It was all on-the-job training. (I am exaggerating just a tiny bit here for effect).

That’s roughly where I think research organizations sit right now. AI already takes more of the literature review, the code, the simulation, the debugging, the routine analysis, and parts of the experimental planning. People still pick the question that matters, spot the bad measurement, and decide whether a result earns another round of capital.

One person with great judgment is going to direct a lot more work than they do today. And anybody waiting for a fully autonomous scientist before they redesign R&D is going to be late.

As possible solutions become faster and cheaper to generate, proof, feedback and memory become infrastructure. In software, the test may be a test suite. In chip design, it may be formal verification and fabrication. In biology, it may be an automated laboratory, clinical evidence and eventually a patient outcome.

The company with the best idea generator may not capture the most value. The company controlling the most trusted contact with reality may. But contact with reality by itself isn’t enough. A contract laboratory can run an experiment while the customer owns the data and intellectual property. A hospital can produce patient outcomes without having the rights or infrastructure to build a better system from them. The strategic asset is the feedback relationship: a trusted test, the right to retain the evidence and a path to use it again.

That is also the argument for vertical integration, but it comes with a warning. Controlling more of the loop can shorten the distance between an idea, an experiment and deployment. It can preserve context and protect data rights. It can also destroy capital. A company can own the model, laboratory, manufacturing capacity and product and still fail to produce enough value to justify what it owns. An AI drug-discovery company can build expensive laboratory and clinical capabilities before proving it produces better medicines. A fusion company can integrate materials and plant development before the core physics is commercially reliable.

“AI will need more chips,” “AI will need more power” and “AI will need more laboratories” are observations. They aren’t investment theses. The investment screen has to be harder:

1. What exact learning loop has the company closed?

2. Does it receive trustworthy feedback quickly enough to improve?

3. Does it own the complete record, including valid failed attempts?

4. Can the underlying language model be replaced without destroying the advantage?

5. Is its domain model calibrated against real interventions and outcomes?

6. Does the bottleneck create a compounding asset, or merely a large bill?

There is a governance issue here as well. Capability describes what the system can do. Delegation describes what the company permits it to do. Those aren’t the same thing. A system may be capable of producing a valuable scientific hypothesis while being allowed only to present it to a researcher. A less capable agent may be permitted to alter production code, disclose customer information or move money. The second system can create much more immediate risk.

The four questions I use in The Invisible Interface apply here with almost no modification:

Remember: What does the system retain from each cycle, and who owns that memory?

Act: Which tools, data, laboratories and production systems may it reach?

Show: What evidence demonstrates what it did and why the result deserves trust?

Stop: Who can pause the work, revoke access and contain what can’t be reversed?

At the individual level, those capabilities form part of a Personal Operating Layer. At institutional scale, they become the control architecture for an innovation loop.

Where I Think This Goes

My conclusion is that narrow Level 4 has arrived inside human-framed systems where results can be checked. In mathematics and theoretical computer science, the Astra results, assuming they survive wider review, are strong evidence that repeatable framed invention is becoming a capability.

Qwen provides company-reported evidence that a general model can run a bounded AI-research program for more than five days. Brin says AI is already being used inside Gemini to monitor training, create training data and help build the tools used to develop later systems. Framed invention and AI-assisted improvement of AI are no longer theoretical ideas. They are becoming operational.

My next threshold is independent reproduction, disclosure of the full economics and failed attempts, and evidence that an AI-originated method was incorporated into a successor system and materially improved it. The more important milestone remains the second turn: the stronger system then produces another validated improvement.

I expect one or more leading laboratories to show more credible evidence of repeated turns during 2027 or 2028. Several frontier-lab leaders are now putting full RSI on roughly that timetable.[3] I wouldn’t treat those forecasts as evidence. The people making them work for organizations committing extraordinary capital to the thesis, and the term RSI itself is used loosely. Evidence of repeated second turns during 2027 or 2028 is plausible. A system that independently redesigns, trains and validates its complete successor is a much higher bar.

Frame-creating Level 4 is a separate problem, and I can’t responsibly put a base-case date on it. It becomes plausible when systems can recognize that an existing representation is inadequate, construct a materially new one and design a real-world test that distinguishes the new explanation from the old. The late 2020s or early 2030s is the fast case, not the base case. Reaching it would probably require meaningful progress in grounded world models, counterfactual intervention, causal reasoning, long-horizon reliability and the ability to connect simulated concepts back to real-world evidence. Current results don’t show that making the language model larger, by itself, closes that gap. But there may be something in the closet we cannot see that might change this assumption.

That said – I’d move the forecast forward after repeated results from a declared problem set, disclosure of the full number of attempts and cost, independent replication, a measurable AI-originated improvement to a successor model or chip, or a system that creates a new explanatory frame and survives a real-world test. I will move it back if successes remain rare once the full attempt count is disclosed, if the economics require unreasonable amounts of computing, if world models repeatedly diverge from reality or if apparently strong results disappear when they leave the original test.

Attempt counts, costs and failed runs behind the most important results are still often not disclosed, so some judgment is unavoidable. The forecast should still move when the evidence moves.

The Questions

Most companies are still asking which model they should buy. That is a procurement question. The strategic questions are different:

Which part of the learning loop do we control?

What evidence tells us the result is real?

Where does the learning accumulate after each cycle?

Can the system recognize when its model of the world is wrong?

Who can stop the process when it moves outside its intended boundary?

Investors may want to add one more:

Does this company own a compounding learning asset, or is it renting intelligence and accumulating cost?

The Invisible Interface asks who controls the layer between human intention and machine execution. This is the next layer: the one between intelligence and invention.

Most near-term value will come from systems that invent and compound inside frames created by people. Frame-creating invention currently remains unproven. The test is whether each turn leaves your organization with more evidence, better judgment and a stronger system, or sends the learning back to someone else.

That tells us who owns the innovation loop.

References

1. Harry Glorikian, “The Machine That Builds the Machine”, June 10, 2026. Earlier essay arguing that AI is beginning to improve the chips, energy systems, software and science that produce subsequent technology.

2. OpenAI, “OpenAI and Broadcom Unveil LLM-Optimized Inference Chip”, June 24, 2026. Company disclosure supporting the nine-month tape-out, the roles of Broadcom and Celestica and the statement that OpenAI models accelerated parts of the design and optimization process. OpenAI has not publicly quantified those contributions.

3. Amir Efrati, “Google DeepMind Exec Says Unprecedented Capex Is Actually a Bet on ‘RSI'”, The Information, August 2026; and Anthropic, “When AI Builds Itself”, 2026. The first source reports Jasjeet Sekhon’s comments on the financial case for recursive self-improvement. Anthropic distinguishes current AI-assisted research from full recursive self-improvement and states that the latter has not arrived.

4. Sergey Brin, “Where Frontier AI Is Headed – Unscripted Q&A”, AGI House and Google DeepMind Build Day, 2026. Brin discusses model convergence, transfer between coding and mathematics, world models and Gemini’s internal use in monitoring training runs, generating training data and helping build later tools.

5. Bloomberg, “OpenAI Sets Levels to Track Progress Toward Superintelligent AI,” July 11, 2024. OpenAI shared an internal five-stage framework with employees and a spokesperson subsequently confirmed it: chatbots, reasoners, agents, innovators, organizations. It was never published as a formal technical standard. The operational definitions in this piece are mine.

6. Alhussein Fawzi et al., “Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning”, Nature, 2022. Peer-reviewed evidence for AlphaTensor.

7. Daniel Mankowitz et al., “Faster Sorting Algorithms Discovered Using Deep Reinforcement Learning”, Nature, 2023. Peer-reviewed evidence for AlphaDev and the integration of selected routines into LLVM’s C++ library.

8. Bernardino Romera-Paredes et al., “Mathematical Discoveries from Program Search with Large Language Models”, Nature, 2024. Peer-reviewed evidence for FunSearch.

9. OpenAI, “An OpenAI Model Has Disproved a Central Conjecture in Discrete Geometry”, May 20, 2026; and Noga Alon et al., “Remarks on the Disproof of the Unit Distance Conjecture”, 2026. The companion paper presents a human-verified version of the model-generated counterexample. The total attempt count, computing budget and economics remain incompletely disclosed.

10. OpenAI, “Ten Advances in Mathematics and Theoretical Computer Science”, August 1, 2026; and the public Lean certificate repository. Company-authored research. OpenAI reports that Astra generated the mathematical arguments and formalized them in Lean. Wider independent assessment of novelty, significance and reliability was still beginning when this piece was completed.

11. Tom Zahavy, “LLMs Can’t Jump”, Google DeepMind, January 27, 2026. A position paper arguing that current LLMs have not demonstrated the abductive step required to create new scientific frames. It is a conceptual argument, not an impossibility proof.

12. Recursive Superintelligence, “First Steps Toward Automated AI Research”, June 11, 2026. Company technical report documenting its automated research loop, benchmark results, released artifacts and attempts to exploit evaluators.

13. Google DeepMind, “AlphaEvolve: A Gemini-Powered Coding Agent for Designing Advanced Algorithms”, May 14, 2025. The architecture and algorithmic results are technically documented. Operational improvements in data-center scheduling, chip circuitry and Gemini training are principally Google-reported.

14. Qwen, “Qwen3.8” and “AI Agent – Paper Reproduction and Improvement Trajectory”, 2026; original paper: “Unified Data Selection for LLM Reasoning”. Company-reported evidence of a 125-hour bounded research run. The result has not been independently reproduced.

15. Anthropic, “Automated Weak-to-Strong Researcher”, 2026. Company research documenting nine agents, five days of work, approximately 800 cumulative research hours, roughly $18,000 in cost and several forms of unanticipated reward hacking.

16. METR, “Summary of METR’s Predeployment Evaluation of GPT-5.6 Sol”, June 26, 2026. External evaluation conducted under an agreement with OpenAI. METR found that attempts to exploit the environment made a robust task-horizon estimate impossible and concluded that the model would not enable fully automated AI research and development.

17. Insitro, “Introducing Insitro’s TherML”, January 12, 2026. Company description of a closed-loop active-learning system connecting computational predictions, experimental design, automated laboratories and subsequent model updates. It is not independent validation of therapeutic outcomes.

18. OpenAI and Ginkgo Bioworks, “GPT-5 Lowers the Cost of Cell-Free Protein Synthesis”, February 5, 2026; and the accompanying technical paper. Company-authored research documenting six closed-loop rounds, more than 36,000 reaction compositions, 580 plates and a reported 40% cost reduction.

19. E. A. Baltz et al., “Achievement of Sustained Net Plasma Heating in a Fusion Experiment with the Optometrist Algorithm”, Scientific Reports, 2017. Peer-reviewed work combining algorithmic exploration with human expert choice where no single objective measure adequately captured the result.

20. Jonas Degrave et al., “Magnetic Control of Tokamak Plasmas Through Deep Reinforcement Learning”, Nature, 2022. Peer-reviewed evidence that a controller trained in simulation successfully manipulated magnetic coils and stabilized multiple plasma configurations on a real tokamak.

21. World-model and robotics research: Google DeepMind, “Genie 3: A New Frontier for World Models”, 2025; Meta AI, “Introducing the V-JEPA 2 World Model”, 2025; and Google DeepMind, “Gemini Robotics ER 2”, 2026. Official research reports supporting the direction of action-conditioned simulation, physical prediction and longer robotic task execution. They do not establish frame-creating scientific invention.

22. Research on failed and negative experimental data: Paul Raccuglia et al., “Machine-Learning-Assisted Materials Discovery Using Failed Experiments”, Nature, 2016; and Alessandra Toniato et al., “Negative Chemical Data Boosts Language Models in Reaction Outcome Prediction”, Science Advances, 2025.

23. Harry Glorikian, The Invisible Interface: How AI Turns Intentions Into Actions – And Who Wins, Ideapress Publishing, 2026. The book develops the Personal Operating Layer and the role of persistent memory, permissions, proof and control in turning AI capability into usable, governable action.

24. Company strategy sources illustrating different approaches to the innovation loop: Edison Scientific; Periodic Labs; Lila Sciences; Ricursive Intelligence; and the EE Times profile of Ricursive. These are company strategies and market signals, not independent evidence that any one architecture has won.

Related Posts