A significant shift is emerging in the AI industry that warrants close attention from the capital markets.
Over the past two years, market discussions on AI have centered on several key questions: Can model capabilities continue to improve? Will inference costs decline? Will a trillion-dollar application emerge beyond coding? And will data center construction lead to overcapacity?
However, with Alphabet's release of Gemini 3.8 Flash, SSI's announcement of a 10x compute expansion over the next 12 months, and OpenAI's public discussions on Recursive Self-Improvement (RSI) and training resource allocation, frontier labs are shifting the competitive focus to a new direction.
AI is no longer just serving users; it is also beginning to participate in training, evaluating, and optimizing the next generation of AI.
RSI, or Recursive Self-Improvement, refers to a process where models contribute to more stages of model development: generating algorithms, writing code, designing experiments, invoking training tools, evaluating results, fixing errors, and feeding successful experiences back into the next round of research. If this loop continuously shortens, the pace of AI development could fundamentally change.
And RSI is not a bet by just one company. Alphabet, OpenAI, SSI, Anthropic — nearly all frontier labs are moving in the same direction.
Gemini 3.8 Flash Signals Alphabet's RSI Ambitions
On September 2, Alphabet officially released Gemini 3.8 Flash, its third Flash model in six weeks. The performance metrics alone are remarkable: in independent benchmark tests by Artificial Analysis under high-reasoning mode, the 3.8 Flash scored 59 on the intelligence index, approaching top-tier flagship models like GPT-5.6 Sol and Grok 4.6, at a single-task cost of just $0.58, with an input price of $0.75 per million tokens.
But more important is the characterization from Alphabet DeepMind researcher Shunyu Yao: "For the model, this is just a small step; but for RSI, this is a huge leap." This statement reveals what Alphabet is genuinely pursuing.
Internal memos explicitly state the goal is to "force recursive self-evolution, transform coding models into fully automated AI researchers, and close the entire R&D loop." To this end, Alphabet has assembled a code task force led by the DeepMind CTO with direct oversight from co-founder Sergey Brin, and recruited Barret Zoph, OpenAI's former head of post-training, as VP of Research to oversee reinforcement learning and post-training. The core upgrade of 3.8 Flash is the product embodiment of this strategy: the model is designed to execute more reasoning steps on complex tasks, repeatedly invoke tools, and evaluate and optimize its own results through long-running agent loops. In other words, the model is beginning to take on the full "plan-execute-check-fix" cycle.
Notably, Alphabet had several 3.5 Pro candidate models internally, but all were rejected because they did not offer sufficient advantages over the Flash series. Meanwhile, the next flagship, Gemini 4, remains stuck in post-training. This sequence of events, somewhat serendipitously, has driven the rapid iteration of the Flash series.
OpenAI's Pause and SSI's Acceleration: Giants Moving in Unison
Around the same time Alphabet released its new model, OpenAI CEO Sam Altman made a rare public statement in an interview. He acknowledged that OpenAI had delayed a frontier reinforcement learning (RL) training run — the first time in its history that it proactively paused frontier training. The reason was not a single "smoking gun" event; during training, the team observed the sheer speed of capability improvement. Altman said: "The rate of capability progress... I can only describe it as 'awe-inspiring.' We need more time for safety, alignment, and safeguards to catch up."
He added: "A year ago, I did not believe superintelligence would arrive soon. Now I think it might happen. I am not saying it definitely will, but we are moving at an extremely fast pace." Altman clarified that this pause was specifically for the frontier RL training run, not all training, and clusters were not idle. He emphasized that OpenAI's commercial momentum remains strong, with enterprise revenue now exceeding consumer revenue, and existing models still have significant commercial value to extract.
But the signal of this statement goes far beyond commercial implications: even OpenAI is hitting the brakes to let safety catch up, which is the most direct public signal that RSI is approaching.
Additionally, NVIDIA recently announced a major strategic investment in SSI (Safe Superintelligence Inc.), which simultaneously announced a 10x expansion of compute over the next 12 months. SSI was founded in June 2024 by Ilya Sutskever, co-founder and former chief scientist of OpenAI, with the "sole goal and sole product" of safe superintelligence. The most notable line in the announcement: "Our research is worth scaling."
Market narratives currently suggest that besides coding, AI has yet to find the next trillion-dollar application, implying that AI capital expenditure growth would likely peak around 2028. But the SSI event reveals this narrative may be focusing on the wrong variable. The primary determinant of frontier lab CapEx is never a specific application race, but rather: is the next generation of models still worth scaling up training?
Over the past six months, signals from nearly all frontier labs have been remarkably consistent: OpenAI continues expanding training clusters, Anthropic keeps raising funds for AI infrastructure, xAI keeps expanding Colossus, Meta continues building gigawatt-scale AI campuses, Alphabet expands TPU deployments, and SSI announces a 10x compute expansion. No company's actions indicate that "training is done."
RSI Reshapes Compute Economics: Training Moves From One-Time to Never-Ending
The key may be that RSI is changing the very form of training. The old model was: collect data → train → release → done. The RSI-era model is: Model A generates a new algorithm → trains Model B → Model B optimizes the training process → trains Model C → Model C discovers better RL strategies → the loop continues on. This means training becomes continuous training — it never stops.
To validate a new algorithm, what used to take one training run may now require running 100, 1,000, or 10,000 versions simultaneously, keeping only the best one. This leads to a direct corollary: compute advantage becomes research advantage, and research advantage becomes model advantage. A lab with sufficient GPUs can complete all validation in a day; a lab without them can only complete 5 or 10 per day — the speed gap widens immediately.
OpenAI has put forward the concept of "Automated AI Researcher"; Anthropic extensively involves Claude in model R&D; Alphabet's AlphaEvolve is already using AI to discover new algorithms. These are real implementations of RSI.
Gavin Baker's Warning: Top Labs May Deliberately Sacrifice Revenue to Train
What does this mean for investors? Well-known technology investor Gavin Baker provided a concrete calculation in a recent public dialogue: Suppose a lab possesses 10GW of compute, with 8GW allocated to inference. At $60 billion per GW, that yields annualized revenue of approximately $480 billion. "Assume they achieve a major research breakthrough and decide to compress inference allocation from 8GW to 2GW, while expanding training from 2GW to 8GW. Annualized revenue would then drop from $480 billion to $120 billion. I believe they genuinely might make such a decision, and this is something the public markets must get used to."
Baker also stated that based on firm belief in scaling laws, top model companies will not pursue free cash flow in the near term; instead, they will reinvest all profits and capital into purchasing compute and model training. He noted that this differs from internet companies like Meta and Alphabet — their fundamentals are relatively stable, and there is no massive trade-off between infrastructure costs and revenue. Frontier labs, however, may at any time deliberately sacrifice significant short-term commercial revenue for long-term technological advantage. This is a major risk signal for AI stock valuations.
The Researcher Mindset: The Powerlessness of Compute Dominance
This race is also changing the psychological state of frontier researchers. Sarah Guo, founder of Conviction and AI investor, shared observations from interviews with approximately 250 frontier AI builders in a video blog interview. Over the past 12 months, a growing number of top researchers have come to believe that once an AI research model with recursive self-improvement capabilities arrives, humanity could be only one to two years away from exponential intelligence.
But at the same time, as model training budgets approach tens of billions or even hundreds of billions of dollars and teams expand to thousands of people, some top researchers have developed two negative emotions: "What I do doesn't matter, because the model can soon do it itself," and "Only the scale of compute matters; individual contributions are diluted." This psychological shift itself is indirect confirmation that RSI logic is being widely accepted within the industry — when researchers start to feel that "only compute matters," it precisely shows how tightly compute and research capability are now bound together.
As RSI Approaches, Market Narratives Need Updating
A conclusion with direct implications for investors is emerging. The market narrative of "AI capital expenditure peaking in 2028" is built on the premise that CapEx is solely driven by commercial application demand. But RSI offers a different framework: what determines future CapEx is not just inference demand and application revenue, but also whether model capabilities still merit sustained expansion of training.
As long as frontier labs continue to believe "our research is worth scaling," training investment will not stop. Based on all public signals currently available, that premise remains unshaken.
Will RSI become the core variable for the next round of AI capital expenditure? It is still too early to conclude. RSI has not yet been proven to reliably deliver exponential capability leaps. Improvements on benchmarks do not necessarily mean models can autonomously continue improving themselves in real-world R&D environments. Higher agent capability also brings higher token consumption, more complex system engineering, and greater safety risks.
But judging by the signals from Alphabet, SSI, and OpenAI, frontier labs are clearly no longer treating RSI as a distant theoretical concept. Alphabet is trying to have low-cost models complete longer, more complex, and more verifiable agent tasks; SSI is validating whether "research is worth scaling" with a 10x compute expansion; OpenAI is discussing how to reallocate resources among safety, commercialization, and frontier training.
These changes all point to a common trend: competition in the next phase of the large model industry may not just be "whose models answer questions better," but more importantly "who can make AI participate in AI R&D faster, and complete self-iteration within safety boundaries."
If this trend continues, the core questions for the AI industry chain will also shift. The market should not only ask: Is AI application revenue sufficient to support data center construction? It should also ask: Is the capability curve of frontier models still trending upward? Can AI genuinely shorten R&D cycles? Can training clusters convert compute into research output? And can safety systems keep pace with the speed of model self-iteration?
This, perhaps, is the most significant impact RSI will have on AI capital expenditure, model competition, and technology investment.