Just 18 months after its founding, AfterQuery is on track to become the fastest-growing unicorn in Y Combinator's history. The artificial intelligence training data company is raising a new funding round at a $3.2 billion valuation, according to sources cited by Forbes. The round has a confirmed lead investor, but the amount raised and investor identities have not been disclosed, and AfterQuery declined to comment. YC partner Gustaf Alströmer stated that if the deal closes under current terms, it would break the record for the fastest time from founding to unicorn status among YC companies.
Just five months ago, AfterQuery completed a $30 million Series A round led by Altos Ventures at a post-money valuation of $300 million. The new valuation represents a more than tenfold increase. The AI training data sector has already produced a series of young founders who achieved overnight financial success. Scale AI founder Alexandr Wang became a billionaire at age 24, and the three founders of Mercor joined the billionaire ranks at age 22 following a multi-billion-dollar valuation. Now, AfterQuery has replicated this path at an even faster pace.
The question remains: why can this sector, which no longer sounds fresh or exciting and is already crowded with high-value companies, still accommodate new unicorns? In reality, AfterQuery did not initially plan to sell training data. When founder Mateega and his high school friend Carlos Georgescu entered the YC winter program in early 2025, they intended to develop AI agents for financial work. During testing, they discovered that existing models could generate seemingly complete answers but frequently failed to handle the assumptions, exceptions, and trade-offs inherent in real financial tasks.
The team ultimately abandoned the application-layer product and pivoted to collecting the training materials models were missing: how financial analysts build assumptions, how lawyers revise contracts, how doctors rule out diagnoses, and how software engineers fix issues in existing code. AfterQuery processes these workflows into supervised fine-tuning data, evaluation sets, scoring rubrics, and reinforcement learning environments. Compared to traditional labeling tasks, experts deliver not just a single correct answer but also the sequence of operations, decision rationale, tool usage, and error correction. The company claims its expert network now includes nearly 100,000 practitioners across finance, law, healthcare, and software engineering. Public customers include Nvidia, legal AI company Legora, and Thinking Machines Lab founded by Mira Murati; some of the data has been used in Nvidia's Nemotron models. These customer relationships and expert scale figures come from company disclosures.
This narrative of encoding professional judgment into data is not new. AfterQuery is fast, but it entered a track already occupied by several high-value companies. Scale AI, the most well-known player, initially relied on large numbers of outsourced workers to label images and other basic data. As large models evolved, it expanded into expert data, model evaluation, and post-training. In 2024, Scale AI raised $1 billion at a $13.8 billion valuation; when Meta invested $14.3 billion in 2025, its valuation exceeded $29 billion. Mercor took a more direct approach by organizing professionals into a training data supply network. It recruits doctors, lawyers, programmers, bankers, and researchers to design tasks, answer questions, and evaluate results for model companies. Mercor raised funding at a $10 billion valuation in 2025 and has reported plans to raise approximately $500 million at a $20 billion valuation in 2026.
The high valuations of Scale AI and Mercor demonstrate that capital has long recognized the value of professional human data. AfterQuery's opportunity does not stem from discovering an untouched market, but from the shifting targets of model training. Early data labeling primarily taught models what an image contains or which response was better. Reasoning models and agents now need to master longer task chains: which software to open, which files to read, how to make trade-offs when encountering conflicting information, how to recover after execution failures, and how to determine whether a task is truly complete. Scale AI stated in February that nearly half of its new data training projects already involve reinforcement learning environments. These environments replicate software and business states, record every step taken by an agent, and use expert-designed verifiers to check actual outcomes rather than merely judging whether the final text appears reasonable.
Low-cost labeling has not disappeared, and premium expert data has existed for years. The new trend in 2026 is a third category of commodity emerging alongside these two: work trajectories accumulated by real enterprises. AfterQuery has reportedly acquired the codebases of approximately 30 to 50 failed startups, while Turing has acquired around 5 to 10. AfterQuery transforms real code into training tasks: models attempt to add features or fix problems, and the code previously committed by engineers serves as a reference to generate post-training data. Turing launched "Project Lazarus," publicly acquiring private codebases and associated materials. In its published pricing, legacy codebases typically range from $10,000 to over $100,000, while modern codebases cap at around $10,000; Jira tickets, product requirement documents, architecture materials, and customer service records can also be priced per item.
Behind this data, the true value lies not in any single piece of code but in the process by which the code was formed. A product deployed in real use leaves behind a trail of requirement changes, design debates, failure records, user complaints, and patches. Together, these reveal how a company made trade-offs between cost, time, safety, and customer demands. Such materials rarely appear in open source projects or on public web pages, and they are difficult to fully replicate with synthetic data.
The bankruptcy auction of Spirit Airlines has pushed this competition into a more public arena. Google bid $10 million for Spirit's internal data, surpassing Mercor's $7.5 million offer; AI data company Micro1 subsequently submitted a late bid of $12.5 million. The materials for sale include approximately 100 million emails, 500 million Teams messages, 30 million lines of code, and selected employee records dating back to 1986. As of September 2, the transaction still awaits a bankruptcy court ruling; Google holds the bidding advantage but has not yet formally acquired the data.
Deals like this continue to expand the boundaries of expert data. Knowledge exists not only in the minds of individual lawyers or engineers but also within the accumulated emails, tickets, code modifications, and collaborative relationships of a company over many years. The public internet preserves a wealth of final outcomes but rarely records the hesitation, rework, and compromises that occurred before those outcomes were achieved. What corporate archives preserve is precisely the part that models find hardest to learn from public corpora.
At the same time, the closer data gets to real work, the harder its privacy and security risks become to manage. After Mercor suffered a data breach this year, Meta suspended its collaboration and OpenAI launched an investigation into the incident. The leaked content may have involved model companies' training methods, contractor information, and proprietary data. Spirit's materials are scheduled to be de-identified before delivery, but employee relationships and cross-system records must be preserved, otherwise the data would lose its training value. The Association of Flight Attendants raised objections: when emails, chats, and files can still be linked along the behavioral trail of the same employee, the risk of re-identification does not completely disappear.
Questions remain unresolved: who owns the data employees generate at work, who has the right to sell it after a company collapses, and whether anonymization can fully cover trade secrets and personal relationships. The closer data companies get to real work, the more valuable the training signals they obtain, and the more concentrated the privacy and intellectual property risks they must bear.
The $3.2 billion valuation once again demonstrates that investors do not believe data services will rapidly depreciate as large models become smarter. Every step models advance toward real work requires training data to move in parallel: from answers to processes, from public corpora to internal corporate records, and from a single expert's judgment to a complete record of organizational operations.