The most common trap companies fall into with AI is treating a flashy demo as proof that a product is ready for real-world use. Chris Churchman, who leads Goldman Sachs' digital platform Marquee and co-chairs the global banking and markets AI working group, recently offered a harsher benchmark on a podcast: a demo only showcases the model's single best performance, but once an institutional product goes live, the worst output is what clients and the market scrutinize. This hits the true fault line between proof-of-concept and production. Connecting frontier models to enterprise data can easily generate fluent, polished, and professional-looking answers, but in high-stakes sectors like finance, "roughly correct" has no value. Whether facts are accurate, calculations can be independently verified, conclusions stay within permissions, and errors can be halted promptly—these are what decide if a system can enter real workflows. The interview outlines four boundaries for productization: never substitute demos for acceptance testing, don't bet on a model's temporary flaws, rework workflows, and never outsource human judgment.
Goldman Sachs' AI Experiment
Marquee is Goldman Sachs' digital platform designed for institutional and corporate clients, aiming to help them make decisions under extreme uncertainty and plug into its investment processes. The platform aggregates millions of research articles, trading desk commentary, pre-trade analysis, and content from dozens of data vendors. Churchman compares its Market View feature to "a Pinterest for capital markets," where every component is crafted by domain experts, and the platform already hosts millions of analytical components. Marquee AI, currently used only internally at Goldman Sachs, attempts to unify these fragmented capabilities into a single entry point. When a user poses a question, the system breaks down the research topic, finds relevant research, trading desk views, and analytical components, writes and executes Python calculations as needed, and returns conclusions backed by evidence. The goal is not to build just another chatbot, but to make complex institutional capabilities accessible through natural language.
Finance Cannot Accept AI That Is Only "Close Enough"
Churchman noted in the interview that even a model with 90% factual accuracy may have no practical value because users don't know where the errors are hidden in the remaining 10%, forcing them to verify everything anyway. AI appears to save writing time, but it shifts the cost to line-by-line validation. Ordinary users might tolerate an occasional inaccurate recommendation, but institutional products must rigorously handle permission errors, stale data, conflicting sources, tool call failures, and abnormal market conditions. Average accuracy rates cannot answer where errors occur, their impact, or whether the system can be safely taken over. Churchman used a vivid analogy: once facts and inferences enter the context window, they both go through the same "sausage machine," and the model cannot distinguish between them—not because it deliberately misleads, but because it simply doesn't know the difference. Therefore, Goldman Sachs set a target for Marquee AI: every significant statement must be traceable to research materials, trading desk commentary, or auditable calculations. This kind of sourcing is not about appending a few links at the end of an answer; it means each judgment maps to specific evidence and the calculation process can be reproduced. For enterprises, acceptance testing of an Agent cannot rely only on standard questions and normal workflows; it must actively include stale documents, contradictory data, insufficient permissions, API timeouts, duplicate submissions, and unanswerable cases. What needs to be tracked is not just accuracy rates, but also citation traceability, human takeover rates, unauthorized actions or errors, failure recovery time, and the total cost of completing a qualified task. The gap between a product and a demo is precisely this unglamorous engineering work that rarely makes it to launch events.
Don't Bet on Model Flaws; Build What Models Can't Learn
Churchman's second point is: don't bet against improvements in model capabilities. He recounted that early models only had roughly 4K tokens of context—about 3,000 English words—forcing developers to chunk data, embed vector databases, and research ways to avoid the "lost in the middle" phenomenon. But context windows later grew to 128K and then one million tokens, long-text retrieval improved rapidly, and function calling and MCP absorbed much of the tool-connection scaffolding. Product moats built around temporary flaws will likely depreciate quickly after the next model release. This doesn't mean RAG, workflow orchestration, or Agent frameworks are obsolete. What changes is that they shouldn't merely compensate for a model's inability to read long texts or use tools; they should handle tasks models cannot solve independently: identifying which source is authoritative, when knowledge expires, what the current user is allowed to see, how different data relates, where conclusions come from, and who must approve write operations. Churchman sums these up as institutional details a model cannot learn from vast public training data, including proprietary enterprise data, permission systems, data connections, and business authorization. For B2B AI suppliers, this means general generation, search, and tool calling will become commoditized rapidly; more durable value will shift toward enterprise context, system integration, permission governance, evaluation, and ongoing operations. Enterprises building their own solutions should follow the same principle: don't pour resources into replicating general capabilities that model vendors will likely internalize; instead, prioritize building their own authoritative knowledge sources, business objects, interfaces, permission mappings, evaluation sets, and accountability systems. Models will upgrade, but institutional facts won't organize themselves, and organizational responsibility will never grow out of parameters.
The Next Stop for Agents in Production: Authorization Engineering
The interview also sketches a clear technical evolution: from prompt engineering, to context engineering, to Agent engineering where the Agent inspects and corrects its own work, and onward to environment engineering and what is not yet a standard term—authorization engineering. Environment engineering addresses where the Agent operates: can it enter a secure environment, access the latest knowledge and correct tools, operate under its real identity and permissions, and have every step be observable, pauseable, and rollback-able? Authorization engineering tackles the harder problem: what is the Agent allowed to do, in whose name does it act, based on whose authorization, and who bears responsibility when things go wrong. The stronger the model, the less these issues can be patched after launch. So-called autonomy should not be simply understood as "executing more steps consecutively," but as a system completing tasks independently within clear responsibility boundaries and voluntarily stopping when it hits those boundaries. For high-risk processes, a practical structure typically looks like: AI handles material identification, information verification, analysis, or initial review; key decisions are confirmed by designated humans; results are written back into existing business systems and enter the current audit trail. This is why many PoCs fail to scale. A demo only needs to prove what a model can do; a product must also prove the organization dares to let it act.
Rethink Tasks Rather Than Automate Old Processes
Churchman frames enterprise AI usage into two approaches: automation and reimagination. Automation finds cognitive bottlenecks in existing workflows—reading, reviewing, data entry—and replaces human effort with AI to deliver the same results faster. It's easy to implement and generates efficiency, but it may also ossify redundancies baked into historical processes. Reimagination starts from "what actually needs to be accomplished," assumes intelligence can be called upon cheaply and elastically, and then redesigns the division of labor between humans and software. Churchman summarizes this shift as: in the past, users learned software; now, software learns users. Marquee AI doesn't push users to master more menus; it understands their research intent and then organizes research, data, tools, and computation. He also acknowledges that the native product form of AI has yet to emerge: the industry is like the early days of television when radio programs were still broadcast over it, or still stuck at the command line before Windows and the mouse arrived. Many of today's products—"a chat box over an old process"—may be transitional forms. Companies evaluating whether an Agent project has value can apply the same distinction. If AI only generates a result and employees must manually move it back into legacy systems, re-verify it, and initiate approvals, it may just be adding an interface. A true closed loop should reduce cross-system handoffs, letting results flow into downstream processes while preserving necessary confirmation and rollback mechanisms.
Human Judgment Is the One Thing That Cannot Be Outsourced
The most substantive layer of this interview is the concern about the degradation of human capabilities. Churchman argues that after technology replaced memory, navigation, and information seeking, generative AI is now taking over argumentation and reasoning. If companies only chase reductions in manual steps, younger employees may skip the training required to form professional judgment. Much institutional knowledge is not written down; it is transmitted through mentorship and hands-on practice, and over-automation could sever that knowledge chain entirely. He likens reasoning ability to physical fitness: people need to go to the gym because they sit in an office for 12 hours a day; the AI era demands a "gym for thinking." He hopes Goldman Sachs newcomers a decade from now will be more like fighter pilots—able to process vast amounts of information and make judgments under uncertainty—rather than bus drivers who just follow fixed routes. This is not about rejecting AI, but about requiring products to enhance human reasoning rather than replace it. Therefore, good enterprise AI should not hide the evidence behind its reasoning; it should help humans judge better: show sources and calculations, clarify facts, assumptions, and inferences, expose conflicts and uncertainty, and leave high-stakes choices to accountable people. AI can expand a person's information bandwidth, but it should not strip away the ability to question its output.
The Real Moonshot for Enterprise AI
Goldman Sachs' experience ultimately points to one conclusion: the moat for enterprise AI is not re-wrapping a general-purpose model, but building a working environment a model cannot learn from the public world. This includes authoritative data and its relationships, identity and permissions, task authorization, tool interfaces, exception escalation, business acceptance criteria, and organizational memory. Model upgrades will consume much of the temporary scaffolding, but they will not automatically answer for a company: which facts can be used, who has the authority to act, what results are acceptable, and when decisions must be handed back to humans. For suppliers, the next phase cannot just showcase what an Agent "can do"; it must also explain how to stop, investigate, and recover in the worst-case scenario. For enterprises, procurement should not start from feature checklists; instead, select real tasks and test completion rates, traceability, human intervention, and unit costs under actual permissions and abnormal conditions. A demo proves AI has the capability; a product proves the organization can trust that capability. Moving from the former to the latter is where B2B AI applications truly begin to create value.