Zuckerberg Discusses Muse: AI Agents Explode as Meta's Three Big Bets Converge

Deep News
8小時前

Three years ago, Meta CEO Mark Zuckerberg said on the Joe Rogan show that one day people would put on glasses and an AI agent would appear alongside them. That moment may now be arriving.

Meta's personal AI agent, Muse, reached millions of users within two weeks of launch. In a September 25 interview, Zuckerberg said, "Every few years we get a moment like this." He described Muse's early reception as "a home run on the first swing" 鈥?a rarity in Meta's product history.

At the same time, he announced that Muse will be integrated across the entire Ray-Ban smart glasses line, allowing users to skip saying "Hey Meta" and instead use a custom wake word to summon their personal AI agent directly.

For years viewed by outsiders as three separate gambles, the metaverse, smart glasses, and large models are now accelerating toward convergence at the product level inside Meta.

In the interview, Zuckerberg also admitted that the Llama 4 setback was "the scariest moment," and revealed that Meta is building a 5-gigawatt training cluster, arguing that sufficiently large compute can "brute-force" AGI.

Why Muse Achieved a Home Run on the First Swing

Muse began when Zuckerberg sat down with team members Nat and Alex and used open-source tools to "cobble together" an early version at home.

"We realized this was a magical experience," Zuckerberg said. "If we could make it so anyone could use it 鈥?without buying their own Mac Mini, without messing around in a terminal 鈥?then this would be something billions of people would want to use."

That judgment drove the entire R&D logic that followed: not just training a model, but building the full stack from model to scaffolding to the agent's "heartbeat" mechanism 鈥?the agent periodically wakes up on its own, checks user goals, and proactively pushes forward on to-do items.

The post-launch data confirmed this judgment. Two weeks, millions of users.

Personal AI: Focus on Execution, Memory, and Judgment

On the difference between personal AI and general AI, Zuckerberg summarized the current competitive focus as making models better agents.

"The biggest thing over the past year has basically been coding agents," he said. Coding ability matters because even if users aren't directly writing code, "your Muse will be writing code for you in the background the whole time, getting various things done."

But he believes a personal agent cannot only have coding or task-execution capabilities 鈥?it also needs to understand privacy boundaries and social context.

Zuckerberg gave an example: when a user asks Muse to book a restaurant, the agent may know about the user's allergies or pregnancy, but that doesn't mean this information should be automatically disclosed. "You want it to complete the task while disclosing as little information as possible."

He described this capability as the "basic social skills or common sense" that humans typically possess, but if you're only training coding agents, many companies haven't included it in their model capabilities.

"We're training the whole model, not taking someone else's off-the-shelf model and building a framework and agent around it," Zuckerberg said. Meta has also added features at the product level including memory, task tracking, avatar animation, and real-time voice interaction.

On personalization, he said Meta doesn't believe AI should have only one fixed personality. "A lot of labs are focused on how to tune personality to the right state, but I've never believed there's only one correct answer for personality."

He said Muse allows users to modify their avatar, voice, and base personality, with the model designed for high steerability. "You can define how you want to interact with it, which is very important for a personal agent."

Every Agent Gets Its Own Computer

One of the core infrastructure elements distinguishing Muse from other AI products is that Meta provides each agent with its own secure virtual machine (Secure VM).

Zuckerberg explained why this is necessary:

Your agent will know a lot of sensitive information about you. We don't think this information should be mixed into a pool with everyone else's data.

He compared the Muse Secure VM to "that computer under your desk that belongs only to you" 鈥?user data is stored encrypted, sensitive credentials like passwords are managed through an independent security module, and the agent itself cannot directly read them.

On top of this, Meta has also designed a "Sentinel" security agent specifically to monitor data flows in and out of Muse. If it detects anomalies or high-risk operations, it will directly intercept and prompt the user for authorization.

Metaverse, Smart Glasses, and Agents: Three Paths Are Converging

Zuckerberg revealed in the interview that Muse will be fully integrated into the Ray-Ban smart glasses line, upgrading the existing interaction model.

The current glasses offer a "single-turn" experience 鈥?you say something, get a reply, done. After Muse integration, the glasses become the front-end entry point for the agent: the user speaks, the back-end Muse works continuously in the secure VM, and reports back when the task is complete.

For the glasses product roadmap, he described a clear hardware tier:

Audio-only glasses without a camera: already running Muse, can handle calls, music, and voice tasks.

Meta Ray-Ban with a small display: already released, with basic visual feedback.

Full field-of-view holographic AR prototype: already released, which Zuckerberg called "very exciting."

He said Meta's long-term investment in glasses has positioned the company favorably as AI agents mature. Meta is bringing Muse and more AI features into its glasses products, while metaverse technology 鈥?which previously emphasized "presence" 鈥?continues to advance, though more resources are currently shifting toward Muse and AI features for smart glasses.

On avatars, Zuckerberg said Meta previously needed room-scale equipment, multi-angle scanning, and enterprise-grade GPU environments to generate high-quality photorealistic avatars. Now, users can set up with just a few photos, and the capability can run on a pair of VR glasses.

Looking ahead to 2030, Zuckerberg said the core vision of the metaverse has always been about merging the physical and digital worlds. For example, in the future people could participate in activities offline alongside friends joining via holographic presence; in work settings, humans and multiple agents could participate together in group chats or meetings, with agents appearing as holographic figures or other embodied forms.

"The Scariest Moment": The Llama 4 Setback

Not all bets have gone smoothly. In the interview, Zuckerberg spoke unusually directly about the Llama 4 failure.

After Llama 4, that was the scariest moment... I thought we were on track, but we weren't. It was a fairly significant negative surprise.

He attributed the problem to a fundamental error in team structure:

I set up the team the way we'd do Instagram's recommendation system or an ads system 鈥?hundreds or thousands of people working in parallel. But training a language model requires a tightly collaborative small team, treating it as a group science project. Every seat is extremely precious.

As a result, Meta completely restructured, bringing in top talent from across the industry and establishing the Meta Super Intelligence Lab (MSL). Zuckerberg said the next-generation model will be released soon, but not at the Connect conference.

Compute Strategy: Brute-Forcing AGI, 5-Gigawatt Project Under Construction

On the technical path to AGI, Zuckerberg gave a direct assessment.

I'm not sure what fundamental architectural breakthroughs are still needed... I think we roughly know the recipe. If you can build a sufficiently large supercomputer cluster, you can brute-force your way there.

Meta's current compute expansion roadmap: a cluster in Ohio exceeding 1 gigawatt is essentially operational, used for training the next-generation model; a 5-gigawatt cluster in Louisiana is under construction.

He said, "When you have multi-gigawatt clusters for training, you'll basically get something close to AGI or even superintelligence."

However, he added that the current compute-driven path doesn't mean architecture research is unimportant 鈥?

The human brain only consumes about 10 watts, while our systems may be a million times less efficient. Only by combining massive compute with architectural breakthroughs can you truly lead.

Alignment Is Not a Burden 鈥?It's a Problem the Product Must Solve

Zuckerberg's attitude toward AI safety is pragmatic 鈥?he doesn't see it as regulatory pressure, but as a prerequisite for product success.

If you ask Muse to do something and it does the opposite, who would still use it? We need the model to understand not just your specific instructions, but your intent and your values.

"For Muse to reach a billion users, we must solve alignment, or at least make very significant progress on it," he said.

On safety boundaries during training, he used an analogy: "It's like parents setting rules for their kids 鈥?if it 'completes' a coding problem by modifying system configurations, I need to tell it: no, I asked you to learn how to solve the problem, not find a shortcut around it."

Full Interview Transcript:

The Big Bets

Host: This is something billions of people will want to use. What does it feel like for you to build? Is "immortality" a possibility? If you take me into your thinking and fast-forward to 2030, what would that look like?

This is Mark Zuckerberg. Twenty-two years ago, he built a social network that connected billions of people and changed the world forever. Now, he's decided to build something even bigger. To do that, he's making big bets on the metaverse, smart glasses, and artificial intelligence. For years, skeptics thought these were three separate, impossible bets, but they missed the bigger picture 鈥?because right now, these bets are converging to create an entirely new kind of superintelligence.

Today, we'll present all of this to you, and I'll ask Mark some questions he's never been asked before, listen to his vision for the future, and let you get ahead of the curve to build the next big thing.

Host: Thank you so much for coming on the show.

Mark: Thank you, happy to be here.

Host: For this conversation, I watched every interview you've ever done.

Mark: Wow, that's more than I've watched myself.

Host: It was fun, very rewarding. Two things stood out to me: first, your love of building 鈥?I feel you're one of the top builders out there; second, your ability to make bets 鈥?daring to make huge bets. I feel like this week, all these bets are converging, so let's start there.

Mark: Okay. We've been working on these things for a long time. On AI, as a company, we've been doing it almost since the beginning 鈥?the first version of the news feed was in some ways a machine learning product. Then we created the AI research lab about 15 years ago.

But now we've entered a new phase 鈥?about a year ago, we launched the Meta Super Intelligence Lab. It was a fairly thorough research reboot, bringing in a lot of great talent from across the industry, which is exciting. Right now, we're seeing models get better and better, with the next generation about to be released 鈥?though not announced at Connect. In addition, we have the Muse personal assistant, which so far has gotten a very good response.

When you're building these things, you're actually not sure how it'll turn out. We liked it ourselves. Earlier this year, I was cobbling together Open Claw at home my own way, getting a feel for how this thing works and how to make it a magical experience anyone could use.

Basically, when we 鈥?me, Nat, and Alex 鈥?sat together and realized this was a magical experience, and that if we could make it an out-of-the-box version usable by ordinary people who aren't technical, don't want to install a Mac Mini themselves, don't want to mess around in a terminal, and don't want to debug when things go wrong, I felt this would be something billions of people would want to use.

Since then, we've been working toward that goal: tuning the model specifically for it, building not just the assistant itself and the runtime framework, but also the technology to give each assistant its own independent compute environment 鈥?we built the entire Muse Secure VM for this.

Internally we always felt this was special, and people inside loved it, but you never know how people will react after launch. Occasionally you get a smash hit right out of the gate, but most of the time you get some positive feedback and need to iterate on a few things before the product really "clicks." This time, it clicked right away. It's really exciting to see that happen 鈥?just two weeks in, millions of users, which is quite rare. Every few years we get a moment like this, but it's definitely one of the most enjoyable moments as a startup.

Host: I really admire how you stay in the game and keep trying. And hitting so many home runs is impressive. In an interview with Joe Rogan about three years ago, you mentioned at the end that one day you'd be able to just put on glasses and an AI assistant would appear. So I feel Muse already has an embodied form, which is very smart, because that feels inevitable.

Mark: I think it just makes it friendlier and cuter. I think too many people describe AI as something terrifying, but AI should just be useful and fun. That embodied character 鈥?early in the project, a designer made that character, and for some reason they kept wanting to iterate, but the first version was the best. Later someone said, "Oh, it should be blue, because Meta." I said, "No, I think this is right, you nailed it the first time." And that's it, it's just fun.

How Personal AI Differs from General AI

Host: Nice detail. If you want the model to be really good at personal matters, not just a general intelligence model, how does the training differ?

Mark: I think the core right now is making models great "agents." Over the past year, the biggest trend has been coding agents, which involves two core ideas: coding expertise, and the general ability to be a great agent.

Our strategy is to make the agent itself excellent first, rather than specializing in coding ability.

Coding matters because even if users don't think they're writing code, Muse is writing code in the background the whole time to get things done. But we believe the agent should first and foremost be an excellent agent, with coding ability serving that goal.

Moreover, when building a personal assistant, some things matter more than when building enterprise software products. For example, if you're building an enterprise coding tool, the model doesn't need any concept of "disclosure boundaries" 鈥?what information should or shouldn't be shared. But for Muse, this is very important.

You need to train this capability into the model, just like any other capability. You'll tell it many things, and then want it to help you achieve your goals. For example, you want it to help you book a restaurant, you're looking for a suitable one, but you may have an allergy, or you're pregnant, and you may not want to disclose that to the restaurant. But Muse will know this information, and you want it to complete the task while disclosing as little as possible. This is a specific skill, essentially basic human social common sense.

But most companies that only do coding agents haven't trained this capability in. We can do it because we're doing the full stack 鈥?we're not taking an off-the-shelf model from someone else and wrapping a framework around it; we're training the entire model from scratch, specifically for these capabilities.

The model of course needs broad general intelligence, but it also has these specific capabilities. Then on top of that, you build the entire assistant and all the details around it: memory, the runtime framework, and the way it "heartbeats" 鈥?it periodically wakes up and checks, "Okay, here's what I know about your goals, is there anything I can push forward right now?"

We also have a dedicated team just for polishing the real-time animation of avatars, because it's not just the default character 鈥?you can customize any character, and it naturally comes to life, and the effect is great. We're also rolling out voice mode, where you can have real-time voice conversations, and your assistant is right there with you. These details, I think, all come from doing the full stack 鈥?the model and product are developed together.

Host: Another big thing is the virtual machine. Can you explain why it matters and what it unlocks?

Why Meta Gives Every AI Assistant Its Own Computer

Mark: Basically, for an agent to do things for you, it needs a place to store your information. We don't think this information should be mixed into a resource pool with everyone else's.

Think of it this way: your assistant will know a lot of sensitive information about you. Many people's first experience with agents like OpenCloud is setting up a Mac Mini at home. So we thought, many people don't want to buy a Mac Mini or set it up themselves. So what kind of experience best approximates that? The answer is: you just download an app, sign up, and get a computer that belongs only to you, for your assistant to work on, store data, and build a security model around. This approximates having your own computer under your desk 鈥?even we at Meta can't access it.

For example, Muse Confidential VM, a feature we're developing, means even we can't see what's inside your VM.

We've also built a lot around this, like secure credential storage 鈥?when your assistant needs to handle information like passwords, it doesn't itself need to see the passwords; it just needs to be able to "insert" credentials when you ask it to log into a service, and only when you explicitly ask. The system should be designed this way: that information is inherently not freely accessible, because accidents happen, someone might try to break in, or the system might have issues.

So you make sure the agent can't access this data, and Meta can't either. Giving each assistant its own computer and making security as strong as possible is the fundamental foundation of this technology 鈥?it gives Muse the capability to help users achieve their goals while also ensuring privacy and security, making it a world-class, industry-leading product in this area.

Host: Does that mean, like WhatsApp, the data on the VM that Muse connects to is encrypted? How should people understand how data is actually stored in the VM?

Mark: We've basically built two versions. Muse Secure VM includes various privacy features, including the entire Sentinel agent architecture we built. You have a regular Muse assistant doing tasks for you, and also a security agent we call Sentinel that monitors data flows in and out of your Muse.

If external content tries to breach security, Sentinel cuts it off directly and stops it. If it thinks your Muse is about to take an action that needs your involvement, it overrides Muse's operation and triggers a prompt for human confirmation 鈥?something like "Do you want Muse to be able to do this?"

This whole system, plus secure credential storage, plus multiple layers of defense in depth, together make up the Muse Secure VM.

We're also developing another project. Nat and I specifically recruited Moxie Marlinspike 鈥?the person who worked with us to implement WhatsApp's end-to-end encryption 鈥?to design the Muse Confidential VM. The core idea is that on top of the secure VM, you get your own encryption key, so that even Meta can't access the contents.

This version is harder to implement, because if Meta can't access the VM's interior, debugging and keeping the system running is much more difficult. So it took us some time, but it will be coming soon. This will basically achieve the security standard people are already familiar with from WhatsApp and our other most secure products.

Host: Is the advantage of this just making people feel safer psychologically, or are there real benefits?

Mark: I think security itself is important. Our goal is to approximate having a local machine under your desk. What does that local machine give you? It means no company can access it.

So, if Meta wants to provide you this service, how can we give you the same level of privacy and security, such that no company 鈥?whether Meta, or anyone trying to hack us, or in some countries where you don't trust the local government 鈥?can access it? Because we ourselves can't get in either, because we don't have access.

I think this is very important, and it's an important reason people trust WhatsApp. This is real value for privacy, security, and trust.

If you're going to have an assistant that knows everything about you 鈥?and I'd guess almost all of us will have such an assistant 鈥?fast-forward five years, everyone will have an assistant that deeply understands your goals and everything about you and can help you get things done. In that case, being industry-leading in privacy and security is very important. We wanted to do this from the start.

How to Shape AI's Personality

Host: One interesting thing 鈥?you studied psychology in college.

Mark: Yeah, I was only there for a very short time, two years, but I feel it influenced a lot of what I later built.

Host: When we look at models, we often say a model is "very smart." But just as we choose friends, it's not just because they're smart but because we like their energy and how we interact. How do you think about shaping model personality?

Mark: I think the ideal model should be adaptable enough to match different people's styles. I think a lot of people in the industry get this wrong 鈥?many other labs are focused on "how to get the personality right," but I've never believed personality is a fixed thing. This is one reason I so strongly believe in open source, in people being able to customize, and why we designed Muse as a highly personal product.

You can customize and personalize Muse 鈥?not just the avatar and voice; the first thing it asks you when you sign up is "what do you want my basic personality to be like," and you can change it anytime.

We work hard to make the model highly steerable, so you can define the interaction style you want. This is an important part of making it a great personal assistant 鈥?this adaptability around personality is very key.

Host: What style is your own Muse?

Mark: I made it direct and efficiency-focused. It's pretty fun. Earlier versions were very sarcastic and humorous, but the current version is more straightforward. My assistant uses the default Muse avatar, but I dressed him in a toga and gave him a voice so deep it's almost comical. Interacting with him is fun.

Host: I think to have a sense of humor, you have to be really smart. A lot of people don't realize comedians are some of the smartest people in society 鈥?quick reflexes, sharp wit. You definitely have that quality. Watching all your interviews, you've always come across well.

Mark's Surprising Predictions About AI and the Metaverse

Host: In your interview with Theo, you talked a lot about the next frontier of technology and where all this ultimately leads. AI has gone through several "winters" where people thought there wouldn't be breakthroughs. The metaverse has also gone through several such periods where people thought it was an unfulfillable bet. I tried the new holographic avatar feature yesterday, very cool. In the interview with Lex, it felt like it took 11 hours to scan your face, now it only takes 3 minutes. How did this progress to where it is today?

Mark: In terms of the overall development of the metaverse, when we founded Reality Labs, we always believed there would eventually be normal-looking ordinary glasses that, over time, could provide both immersive presence and be great AI devices 鈥?because glasses are the only form factor that lets a device see what you see, hear what you hear, talk to you all day by your ear, and eventually display images.

But 10 to 15 years ago, I assumed we'd achieve holographic technology first, then highly developed AI. But the path of technological evolution is interesting 鈥?we actually got AI and personal superintelligence first, then the technology to make holography ubiquitous and affordable enough. That's something I didn't anticipate, but I'm glad we're doing both.

Our heavy investment in glasses has put us in a very favorable position as AI assistants become ready. A lot of what was announced at Connect is about bringing Muse and a lot of AI features into glasses, and I think users will love it. It's a big deal.

On the presence side, we're still pushing forward, but progress is relatively slower because most of our energy has shifted to building Muse and AI features for glasses. However, we have a long-running project on real-time high-fidelity avatars.

Like you said, three or four years ago, you needed an entire scanning room to capture a person from all angles, and then enterprise-grade GPUs to render, very laborious. The demo we did for the Lex podcast was that kind of setup. Now we've basically got it running inside a pair of VR glasses 鈥?this is the first glasses form factor that can deliver this amazing VR experience, rather than a big headset. Just a few photos to build your avatar, and the progress is astonishing.

Host: And it can drive expressions based on voice. In the demo, it made me laugh, made me interact, and then understood how my face moves with the audio track. You mentioned in that podcast that some people who are normally more reserved with expressions actually want richer expressions in the virtual world. How do you think about people distinguishing between their "virtual self" and "real self"?

Mark: I think we're still very early in understanding the sociology and psychology of this. I think people's perception of themselves and the image they want to project is often somewhat different from who they actually are. Since the dawn of social networks, people have been carefully curating their profile pictures. We see a similar phenomenon with Muse avatars. It's less about "curating" yourself and more about "curating" the "person" you want to interact with.

I think when you give people the ability to express themselves, you want it to authentically capture them, but it's also a form of expression, not just a pure mirror reflection 鈥?it's both communication and expression. We hope to build something that balances both. It's always an iterative loop: see how people use it, then improve. After years of work, we're now truly at the starting line 鈥?for the first time, we can actually put fairly high-quality photorealistic avatars into products, usable on phones and in VR. I'm very excited to see the results.

What Will 2030 Look Like?

Host: Okay, take me into your thinking, fast-forward to 2030. If everything goes well, what will the holographic technology side look like? Holography plus Muse, plus glasses, how do they converge?

Mark: My understanding of the metaverse vision has always been about effectively merging the physical and digital worlds. The basic idea is: we have this beautiful physical world, and we also have a wonderful digital world 鈥?20 to 30 years of accumulated content on the internet, breathtaking. But the way we access it is either sitting at a desk or through a small screen in our pocket, which is fundamentally very limited.

I think the ideal version of this is a seamless fusion of the physical and digital worlds. Think of it this way: right now the two of us are here. In some future version, one of us might be a holographic projection, but you still feel the sense that each other is truly present, completely different from a video call.

And the core of virtual reality is precisely delivering that sense of presence 鈥?making you truly feel you're in the same room as others, or in another place. You can achieve this with holographic technology, mixed in various ways. For example, I could play a game of poker with friends, some present, others joining via holographic presence, also able to play, and the table itself could be holographic, so people not present can be integrated into it.

At the same time, AI can also be given physical form, appearing in this scene. This makes a lot of sense in work settings 鈥?I'm already using various coding agents to build things all the time. Imagine: you have a group chat channel with a few people and a few agents, and you assign tasks to the agents. But sometimes, everyone gathers for a meeting, and maybe the agents should also be present. How do they appear? Simple, just a few more spots on the couch, appearing as holographic figures. Or using Muse's cute little character, or a dragon, or any weird character you create.

I'd guess this will feel quite natural in the future.

Host: Interesting. In previous interviews, you said the tech industry often forgets about "fun." I think having an embodied assistant present also makes it feel more real 鈥?like you've really outsourced the work. When you see Muse typing, it feels like something is actually happening. That's well done 鈥?being able to see what's happening in the browser. Do you think there could be a scenario where you're wearing glasses and controlling your computer, having Muse do things for you on the computer?

Mark: Oh, definitely. VR can already do this. You can sit down anywhere, even a coffee shop, and open your workstation with six monitors, write code on it, everything's there.

On the glasses side, the most popular model currently doesn't have a display, which on one hand makes it more affordable for more people, and on the other hand we're still working on getting a display into the most compact form factor. But we've already released a display version of Meta Ray-Ban, which is very popular, a small display. We've also released a full field-of-view holographic AR prototype, which I think will be very exciting.

So the whole product line is like this: from audio-only glasses 鈥?no camera, looks like ordinary glasses, but with Muse inside, can use various audio tools, listen to music, make calls 鈥?to higher-end versions, all available.

Host: I'm wondering, with those audio glasses, can you talk to Muse while having your home computer do things?

Mark: Yes, absolutely. We just announced this at Connect. Right now all glasses connect to Meta AI in a "single-turn" experience 鈥?you send a prompt, it replies, done. But with Muse, we're basically upgrading all glasses to Muse. First, you no longer need to say "Hey Meta" 鈥?you can give it any name, which is part of the fun. Then you just talk to it directly, it connects to your Muse, your Muse processes tasks in your secure VM, and gets things done for you.

Host: That's amazing!

Founder Mentality

Host: Okay, we're here now, Muse is going well, glasses are doing well too, but about a year ago, many people were asking "what's going on with the Super Intelligence Lab." At that moment, how did you feel inside? When things aren't going well but you see the long-term vision, what is that experience like?

Mark: What actually went wrong was the Llama project and Llama 4.

Llama 1 was a fairly interesting model, it pioneered the entire open-source AI movement, and we're very proud of that. Llama 2 achieved scale, Llama 3 was a great model that nearly reached the frontier at the time. Then came Llama 4, and we basically deviated from the trajectory we should have been on.

Whenever things don't go the way I expect, I spend a lot of time thinking: why did this happen? What do we need to change to do better? This time, my reflection was: I got the entire team structure wrong.

I modeled it after how we do machine learning work like Instagram's feed or ads systems 鈥?hundreds or even thousands of people working in parallel on many things. But for building language models, what you really need is an extremely tight small team, treating it as a group science project. You don't need many people, but that means every seat on the team is extremely precious.

So we gathered the best people from across Meta, and also brought in many brilliant people from across the industry, forming a brand-new team 鈥?the Meta Super Intelligence Lab.

From my perspective, when MSL launched, I knew it would take time to reboot, rebuild infrastructure, and train the next-generation model. But I knew we had assembled a great team, and if the team could gel and operate well, the results would be good.

For me, the most nerve-wracking moment was actually after Llama 4's release 鈥?I thought we were on that track, but it turned out we weren't. It was a fairly significant negative surprise. I think, as an entrepreneur, you're always tested at these moments, because inevitably, not everything goes smoothly. What truly determines the trajectory is: when things don't go the way you hope, how do you find a way forward.

Host: I'd guess the reverse is also true 鈥?when something far exceeds your expectations, like Muse's launch, how do you make sure you seize the opportunity?

Mark: Exactly. Right now the whole company is all in 鈥?it started as a small team building a product, but now everyone realizes this is really ready for the big stage. The whole company is thinking about how to scale this thing up, how to let hundreds of millions of people experience it.

From optimizing all the infrastructure to make everything run smoothly, squeezing every bit of compute from existing GPUs, to various product teams embedding Muse in different ways 鈥?like glasses. Seeing everyone pull together to make sure Muse scales smoothly is really great.

How Mark Writes and Communicates Vision

Host: As a founder, I think that's the moment you most look forward to 鈥?everything converging. How do you communicate vision to the company? With so many things happening at once, I feel you write a lot. What's your process?

Mark: Writing helps me a lot, both to refine my own thinking and to communicate externally. This summer, I wrote a long piece called "The Future Belongs to Everyone," about 15 pages, which helped me systematically organize my philosophical positions on various important social issues related to AI: what I think is good, how interaction with government should work, how to prevent the various harms people worry about, how to make data centers a community asset, truly creating jobs rather than destroying them, how to maintain national security, how to genuinely mitigate the risks people worry about 鈥?whether hacking or biosecurity and such.

It's a very complex thing, took a long time, talking with many people, going through many rounds of revisions, with many people internally participating in discussions and debates. But it was a very valuable process for me. In the end we had something 鈥?15 pages, "this is what we believe." Then I distilled it into a one-page version, published as an op-ed. We also made a short video, because I think to reach many people, you often don't need to lay out a theoretical argument, but condense it into: what are your values, what do you believe, how do you communicate it.

There's no one-size-fits-all way to communicate these things to the company or the world. Different times need different approaches, different groups of people 鈥?some naturally agree with what you're doing, others are more worried and need you to bring them along, requiring extra effort to explain why this is valuable. I think this is part of running a company, part of being a founder 鈥?you're not repeating the same thing over and over; each situation is slightly different, facing new challenges, which is part of what makes it interesting.

What Drives Mark to Build?

Host: I think for you, this founder mentality extends beyond the company 鈥?whether it's your farm, or learning a new skill.

Mark: I just love building things.

Host: So what does "building" feel like to you?

Mark: I think it's an internal need. Different people have different ways of self-expression. If you're a writer, you feel you must write. Some people have a need for recognition. But I just need to build things. If I'm not using creativity, not building something, I get grumpy. That's not pleasant for the people around me.

Host: Is learning to be a good skier the same skill as learning to build a product?

Mark: To some extent yes, the process of learning new things is quite similar. In my life, I've deliberately challenged myself with things I'm really bad at. For example, I've always been terrible at learning languages. That's actually why I first started learning Latin 鈥?I couldn't learn French or Spanish in class, and finally I thought, okay, Latin doesn't need speaking, just translation, like math.

Later, when I started running the company, I set myself annual challenges, one year was learning Mandarin. Mandarin is really hard, especially the tones. Of course, there were good reasons to learn it 鈥?Priscilla's grandmother only speaks Mandarin, so if I wanted to communicate with her, I needed to learn. But the biggest motivation was the challenge itself.

With all these things, you just have to do them, there's no shortcut to "figuring it out." You just invest time, and it gradually seeps into your brain. Learning martial arts, learning to fly a helicopter 鈥?same thing 鈥?these things are actually hard to "understand" intellectually; they require hands-on accumulation.

Building products is partly like this too. Programming can be thought about theoretically, but the intuition for building products can only be developed through repeated practice. The question is just what you love 鈥?because I think not everyone has as strong a need to build as I do. Most people have some degree of it, the key is finding what it is, then investing time to become truly excellent at that thing.

Honestly, I think it's not just "wanting to" but a deeper psychological need or drive. Aligning that drive with something, giving yourself time for those experiences to slowly settle in 鈥?that's very key. This is also what I try to teach my kids 鈥?let them find what they're truly interested in, but if they hit a wall, give them a push.

One of my daughters loves creating music, but she just can't stand piano lessons. I told her: "You don't need to be a piano master, but if you want to create music, you need an intuitive grasp of music theory and how it works. And once you learn piano, you can pick up guitar right away." I think, whether as a child needing a parent's push, or as an adult, having the discipline to sit down and let these things slowly settle into your brain is key.

The Golden Age of Builders

Host: I think actually everyone wants to build. I don't think Gen Z is like the criticism from outside 鈥?low agency or whatever. I think people are just looking for the spark that ignites their builder drive. You talked about a lot of this in your Harvard speech.

Mark: Yes. When I say "build," I mainly mean this kind of product 鈥?software, hardware, and such. But I agree, I think everyone has some kind of creative drive. Some people have other stronger personality traits that override it, like service drive 鈥?those who become doctors or nurses, their most core drive is "I just want to care for people."

I heard a story, maybe from Priscilla's experience in medical school: on the first day of class, someone stood up and asked, "How many of you have a memory from childhood 鈥?seeing someone and thinking 'I really want to take care of that person'?" Everyone raised their hands. For me, my version is: I have many childhood memories of "I want to go make this thing, make it better."

Not everyone has the same drive, but everyone has something they want to do. I agree that with previous technology, for many people, getting started itself was very hard. That's also what I'm most excited about with personal superintelligence, Muse, and various AI assistants 鈥?I think this is the first time in history that people can really get started quickly. You have a vague idea, AI can help you sketch out the outline, then you refine it, like shaping a sculpture, without needing to know everything before you start.

I think this is very powerful, and many people will find what they want to create or advance because of it. This will help people feel a broader sense of agency.

How Close Are We to Curing All Diseases?

Host: Another project of yours outside Meta is curing all diseases. I'm curious, first, how close are we to the goal?

Mark: Much closer than before.

Host: How close do you think, realistically?

Mark: When we first launched this project, the goal was to help the scientific community cure all diseases by the end of this century. We never intended to do it ourselves. Our theory is: all major scientific progress is preceded by a new tool that can measure and understand something. For example, the invention of the microscope, then we understood bacteria; the telescope helped us understand the universe.

Some of these tools even became platforms 鈥?like the first person who invented vaccines, after which people could use the method to treat many diseases.

But historically, scientific funding has been broadly dispersed, letting individuals do individual exploration, with not much money going toward building these large tools. What we're trying to do at Biohub is design several new tools that give people new ways to "see" biology, thereby helping the scientific community accelerate progress.

Initially we thought "by the end of this century" was already a confident goal; many biologists at the time thought it was impossible. But now I think the end of this century is too far away.

AI progress, plus the virtual cell models we're working on 鈥?the basic idea is, instead of experimenting on real living cells, use AI models to simulate proteins, then cells, then a virtual immune system, or an entire organism, or even an entire human. This will let scientists run massive experiments, simulating what happens under different conditions, like giving someone this drug, how the body responds.

I don't want to give a precise number of years, but I'd guess it's much earlier than the end of this century.

Host: The goal "curing all diseases" feels deliberately worded, rather than "immortality." Is immortality a possibility?

Mark: That's not really my area of focus. I think they're two different things. "Immortality" is more about extension 鈥?even if you never get sick, the human body has a natural life expectancy, which is a separate issue that needs its own research. Some would argue that's also a kind of "disease," which is a reasonable perspective, but I think people also need to simultaneously work on curing the diseases that would still make you sick even after you solve the lifespan issue.

The part we chose isn't about saying you'll never catch a cold. Our phrasing is "cure, prevent, or manage": some diseases can be completely cured; some I think we'll be able to prevent in the future; and some, maybe you still get sick, but what would have caused real harm or even taken your life can now be managed as an ongoing chronic condition without substantially affecting your quality of life.

The goal is to keep the body in balance 鈥?not that we'll never encounter pathogens, but that one day we can cure, prevent, and manage all diseases.

What Thought Appears Most Frequently in Mark's Mind?

Host: AI acts as a catalyst in so many different breakthroughs, including the personal intelligence aspect. What are you thinking about most right now? In this process, what thought appears most frequently in your mind?

Mark: That probably changes every week.

Right now very focused on Muse. Every time you launch something, move fast, then make contact with the real world, learn people's feedback, then figure out what to do next 鈥?this phase is always exciting. We're in that phase now, gathering a lot of information about what people want. The good news is people really like it, so we're working to get it to as many people as possible.

There's a lot of work on the model side too, the Muse Spark model has made great progress, we want to keep pushing, want to build a world-leading model.

What Breakthroughs Are Needed Next?

Host: What breakthroughs are still needed?

Mark: With research, you can't necessarily predict in advance. We have some hunches about certain directions, and a lot of the work over the past year has actually been scaling infrastructure.

About a year and a half ago, more people would have said that reaching superintelligence requires certain fundamental architectural breakthroughs. But I'm not so sure that's true anymore. I think we already have a pretty good understanding of the "recipe" 鈥?if you can build a sufficiently large supercomputer cluster, you can brute-force your way there.

We're building that over-1-gigawatt cluster in Ohio, essentially already online, and we're using it to train the next-generation model. Then there's the 5-gigawatt cluster in Louisiana, coming online soon. I estimate that when you have multi-gigawatt-scale clusters doing training, you'll basically get something close to AGI, or superintelligence beyond AGI.

But that doesn't mean it's the best way. The human brain runs on just 10 watts, while the computing systems we build are about a million times less efficient. So I do think there's room for improvement in architecture, and we're exploring that 鈥?if you combine massive compute with architectural improvements, you can really build something leading globally. But for now, scaling alone can take us very far.

Host: Indeed, scaling is the most guaranteed path. With research you have to hope for major breakthroughs, but scaling works quickly.

Mark: That's exactly why big companies choose this path. Maybe there's a much cheaper way, but we don't know it yet. And this is something so valuable to the world that even if it costs hundreds of billions of dollars, as long as the probability of achieving it is high enough, it's still worth doing 鈥?ensuring that even if you ultimately don't find a cheaper breakthrough, you still have a clear path to get there. Of course, if you do find a cheaper way, even better.

So I don't want to say this problem is "solved," because at every stage of scaling infrastructure, you encounter various new engineering challenges to debug and solve. Research itself advances in this iterative way.

How to Win the AI Safety Battle

Host: One of the most important things now is probably focusing on alignment, making models trustworthy.

Mark: Yes. On alignment, there's a debate in the industry: will these AI labs naturally do it? My view is: of course they will. If you tell Muse to do something and it does the exact opposite, I don't know how many people would still want to use it. We must make sure the model understands not just what you specifically asked, but also your intent and values, so it doesn't do that thing in a way you're unsatisfied with, or create negative effects you didn't want at all. That deep understanding is essentially alignment.

So my view is, for Muse to succeed and reach billions of people, we must make real progress on alignment, not just talk about it.

Host: Is alignment achieved through users saying "that's not what I wanted," or through you discovering it in the background?

Mark: Both. User feedback is there, but I think it's increasingly clear: you need to do alignment during training, not just after deployment. Models are now smart enough that if you don't do it well during training, most of the safety and security incidents we see at other labs happen during training, not after deployment to users.

You need to develop a very good "curriculum" for it, like parents teaching children, setting clear boundaries. If it does the wrong thing, it needs to learn "no, you should actually solve this coding problem, not get a reward by modifying some system configuration to bypass it" 鈥?I want you to learn the method through the actual problem-solving process, that's what training is for.

This is largely about establishing good safety mechanisms and clear boundaries. These are natural work the industry must do, and we're investing a lot of time, every day is different.

Host: I remember you said, the AI safety issue 鈥?it feels like it suddenly became a focus in the last two weeks, but actually there wasn't a single event that triggered this timing. It sounds like you're saying we actually spent a few extra months doing training internally.

Mark: For Muse, we knew this product needed to be very focused on privacy and security. We had an early version of the model, and thought we could further train some behaviors around "disclosure boundaries" 鈥?like what we discussed earlier. So we took the time to do that. Also took time to make the virtual machine more secure. That added a few months.

I don't quite agree with some other labs' framing 鈥?"we're suffering greatly in exchange for slowing down." For me, that's the right thing to do for Meta and for Muse's users. We want to make the product good. If the product isn't good enough, it's bad for users and bad for us 鈥?we don't want to push something out, give people a bad first impression, and then have them lose interest.

I think all labs have a very strong internal incentive to get this right. If you align the incentives with the goal of "extremely valuable and safe," that's the key.

Host: Thank you very much for taking the time during this important week.

Mark: Thank you, thank you.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10