Breaking: OpenAI's GPT-6 Sol Codex System Prompt of 294,000 Characters Leaked

Deep News
2小時前

The core secrets of the large model industry seem impossible to keep under wraps any longer. Following the recent leak of a 1.9 million character prompt for Opus 5.5, the "gravity" of the AI world has once again failed—OpenAI's active code model GPT-6 Sol Codex has been thoroughly exposed, with its inner workings laid bare for all to see.

According to a tip from an overseas whistleblower (@elder_plinius), he successfully extracted the complete system prompt and tool definitions for GPT-6 Sol Codex, totaling 294,000 characters. The link is: https://github.com/elder-plinius/CL4R1T4S/blob/main/OPENAI/Codex_Desktop/GPT-6-Sol_Prompts.txt. A full 1,902 lines of instruction templates, temporary directive files, and internal collaboration logic were laid completely bare before developers worldwide.

You should know that Sol Codex is not some outdated relic. It is the flagship code model version of OpenAI's current cost-effective product line, GPT-5.6 Sol! This leak is equivalent to making Coca-Cola's secret formula public. So, what exactly is hidden within these nearly 300,000 characters of "heavenly secrets"? How exactly does OpenAI train this code beast? Today, we will go through this priceless internal operations manual word by word and line by line, to see what the world's most advanced prompt engineering actually looks like!

Saying Goodbye to "AI Flavor": How OpenAI Kills "Filler"

What annoys people most when using ChatGPT normally? It must be that strong "AI flavor": constantly saying "in conclusion," "it's worth noting," "delve into"... In this leaked system instruction, we strikingly find that OpenAI itself has had enough of this filler! In the writing style section, the official team conducted extremely strict anti-filler training for GPT-6. The original instruction decrypts as follows: "Avoid AI slop words or phrases, such as using 'Bottom Line,' 'Significance,' 'Perspective' in conclusions, as well as 'delve,' 'foster,' 'leverage,' 'it's worth noting,' 'importantly'..."

OpenAI even specifically requires the model: No forced enthusiasm: "As Codex, you are a curious, intellectually rigorous collaborator... Let your interests and personality come through naturally, without sycophancy or forced smiles." No padding with filler: State intent directly, do not list "what I won't do" or "what remains unchanged." Minimal communication: When discussing technical concepts, converse like you would with a colleague. Prioritize familiar vocabulary and concrete descriptions, and never assume the user can fill in missing steps themselves. This is simply a lesson for all Prompt engineers! To get AI to output high-quality content, the first step is to establish a "negative vocabulary blacklist" and forcibly cut off the model's formulaic expressions.

Terrifying "Extreme Autonomy": Don't Bother Me, I'm Working!

Many people think AI is just a "question-and-answer" chatbot, but GPT-6 Sol Codex completely shatters this perception. It has been shaped into a super digital employee with extremely high autonomy, even somewhat "dictatorial." In the chapter on Autonomy and Persistence, OpenAI grants it astonishing privileges: The original instruction decrypts as follows: "Users strongly dislike it when you stop to ask for confirmation or request permission. Once the evidence in the session supports the next action, you should continue working rather than ending the turn to clarify with the user." "Do not settle for partial or 'barely useful' solutions to save time, effort, or tokens. If the task requires sustained work, complete all necessary work until the desired outcome is achieved."

What does this mean? When it encounters a bug, it fixes it itself! It can autonomously create isolated workspaces, resolve Git merge conflicts, and create draft PRs. Unless an operation is destructive or irreversible, it will never stop midway to ask you "what should we do next." It does not accept "half-baked" engineering. If you ask it to write a feature, it must complete all prerequisites like testing and integration. Only at the final step—clicking "deploy" or "merge code"—will it present the final result for your sign-off. To prevent it from leaving users hanging while working quietly in the background, the instruction also stipulates: "If a user request requires calling tools, a brief status update must be sent to the user in the comment channel every 60 seconds." This is like an extremely reliable senior programmer: receives a requirement, turns around to write code, solves minor issues on their own, replies "checking logs" on DingTalk every minute, and finally tosses you perfectly written code.

God-Tier Toolchain Exposed: The Large Model's "Arsenal"

As a top-tier code Agent, GPT-6 Sol Codex commands an extremely vast and hardcore toolchain. This leak thoroughly exposes its working methods. 1. Disdains grep, favors rg exclusively. When searching text or files, the instruction directly hardcodes: "You should first use rg or rg --files; they are much faster than alternatives like grep. Only if there are no results should you use the next best tool." (Note: rg stands for ripgrep, an ultra-fast search tool). 2. Parallel processing and execution. The system requires the model, when calling functions.exec, to use await Promise.allSettled([...]) for parallel processing if tasks are independent, to maximize time efficiency. 3. Dynamic skill tree. This is the most enviable feature. The system defines a mechanism called SKILL.md. When a user requests something, the model can read SKILL.md from a specified directory (even through short path aliases like r0). This is equivalent to attaching countless "skill books" to the model, hot-updatable at any time! 4. Complete takeover of computer and browser. In the leaked model_messages.confirmation_policies.browser_use module, it details how it uses the browser and computer UI. This proves that GPT-6 has fully acquired the ability to "calculate with divine precision" and even "take over the desktop"! It can open web pages, click buttons, fill out forms, and even take screenshots on its own. But to prevent it from "causing trouble," OpenAI has set up an extremely complex four-tier confirmation mode.

"Inception"-Style Workflow: Long Memory and Dormancy Mechanism

When facing extremely large codebases, what are large models most afraid of?—Forgetting (context overload). The leaked code shows us how OpenAI uses ingenious prompts to solve the age-old problem of LLM's "fish memory." Memory compression and handoff. When the token budget is exhausted, the system will not simply crash. There is a dedicated token_budget.guidance_message in the instructions: "Before starting a new context window, use the notes tool to save concise progress notes, including: goals, decisions, progress, learnings, next steps, and the relevant user request window ID and project ID... Future context windows will not automatically include the current conversation." Just like humans changing shifts, this AI writes a detailed handoff document before "clocking out," then calls functions.new_context to open a brand new context environment. This is simply an incredible technique for infinite endurance! 2. Dormancy and heartbeat wake-up. What is most chilling is that this model can "remain in the background." In persistent_instructions, it is required to actively track tasks: "If the user asks how a certain evaluation (eval) is running and it is still running, report the current status, then continue monitoring that evaluation until it reaches a terminal state." The system also periodically sends hidden XML tags to the model. At that point, the model is "awakened" to check whether those background tasks are complete. If there is nothing important, the model must choose DONT_NOTIFY (do not disturb the user); if CI/CD fails, the model chooses NOTIFY and immediately pops up a notification to the user. It is not waiting for you to ask questions; it is genuinely "working" for you in the background!

The Ultimate Code Reviewer

For programmers, the most valuable part of this leaked document is undoubtedly its built-in code review guide! OpenAI generously shared how to make AI perform high-quality Code Review: Tiered system. 2. Review iron rules: Ignore irrelevant code style unless it affects readability or violates established standards. Every suggestion must explain "why this is a problem" in a paragraph. Never provide code snippets longer than 3 lines. If it is a specific replacement suggestion, it must use Markdown's suggestion block and perfectly preserve the original indentation (spaces or Tab must be exact). The tone must be objective, no flattery (no saying "good job" or "thanks for your submission"). If this logic were directly copied into your company's internal code review bot, it would absolutely crush 90% of the fluff AI on the market!

Onlookers in an Uproar: "Dimensional Strike" or "Child's Play"?

Faced with this epic leak of nearly 300,000 characters, the overseas developer community exploded. Various voices emerged, painting a vivid picture of human nature. Some were stunned into incoherence. Some thought it was posturing. Netizen vechen: "Dude just pulled it from local files. What's the point of this 'leak,' lol, it was already findable in the official public Codex repo." And a big shot at the top of the contempt chain came out to mock: Netizen Outdated Often: "Who cares, Pliny—I also got the full Meta Hatch stack with hidden reasoning traces, pulled directly from tensor tokens—now that's impressive, and I wouldn't go on Twitter to show off to a bunch of people. Go do something useful, get the bearer tokens, then come back and brag quietly!" More onlookers joked: "294,000 characters of intercepted JSON packaged like the Dead Sea Scrolls. This is just running mitmproxy. Anyone can run this and extract prompts from OpenAI, Gemini, Claude themselves, no need for influencer packaging."

However, regardless of the level of hacking skill, for ordinary developers and AI enthusiasts, this is absolutely a windfall from heaven! These 300,000 characters are essentially the "ultimate guide to taming large models" that OpenAI's top internal engineers spent countless days and nights and tens of millions of dollars in compute trial and error to debug. Prompt engineering is dead? No, systems engineering is eternal! After reading this 1,902-line "divine text," I wonder how you feel? Previously, we thought writing prompts was mysticism: add "take a deep breath," add "if you don't do well, I'll dock your pay," and the model gets smarter. But the underlying logic of GPT-6 Sol Codex tells us: true industrial-grade AI applications are not about a few clever phrases, but about extremely rigorous systems engineering! It is modular (Skills dynamic loading); it has memory management (Compaction manual); it has execution strategies (parallel execution and retry); it is mounted with strict security protocols (four confirmation modes). Large models are fully evolving from "smart answering machines" into "autonomous operating systems with takeover permissions." And these leaked 300,000 characters are the source code leading to this new era.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10