Meta Unveils Muse Spark 1.3 With Major Leaps in Coding and Agent Capabilities

Deep News
2 hours ago

Meta's Superintelligence Labs released Muse Spark 1.3 on September 2, immediately making it available on the terminal coding agent Muse Code and the Meta Model API. This iteration marks the fourth update to the Muse Spark series within five months, following the launch of version 1.2 alongside Muse Code on August 5. The new model keeps its weights proprietary, with existing inference tiers available simultaneously, though the highest max reasoning tier remains gated pending additional safety evaluations.

The performance push targets frontier-level models, with coding and agent automation as the primary battlegrounds. Chief AI Officer Alexander Wang told the media this represents the largest performance jump to date, with coding and agentic automation as the main focus. He noted that 1.3 is "competitive" with Anthropic's Claude Fable 5.1 and that its coding performance is "better than" OpenAI's GPT-5.6 Sol. Mark Zuckerberg wrote on social media that this is "the biggest leap yet in coding and agentic workloads," describing it as offering "frontier performance so cheap it hardly needs tracking."

Compared to the previous 1.2 iteration, internal engineer comparisons reveal tool calls have dropped approximately 20% and token consumption has fallen about 25%, with the model "talking less and avoiding unnecessary back-and-forth turns." The context window remains at 1 million tokens, supporting text, image, and video inputs. Standard API pricing is unchanged at $1.25 per million input tokens and $4.25 per million output tokens, with cache hits at $0.15. A contributor tier offers lower pricing, contingent on allowing Meta to use submitted content for model improvement.

Third-party evaluations have placed the model in the top tier, though the max tier remains restricted to partners. Independent review firm Artificial Analysis scored the public 1.3 (xhigh) tier at 61 on its Intelligence Index, tying with GPT-5.6 Sol (max) and Grok 4.6 (high). The partner-facing 1.3 (max) scored 62, trailing only Claude Fable 5.1 and Claude Opus 5. The firm noted that xhigh has the lowest per-task cost among models scoring above 59, at roughly $0.55, compared to approximately $0.94–0.95 for peers at the same tier. However, because agentic evaluations consume about 57% more input tokens than in 1.2, the per-task cost has risen from roughly $0.40 to $0.55.

On specific benchmarks, Tau3-Bench Banking improved from 35% in 1.2 to 47% in xhigh and 52% in max. Terminal-Bench 2.1 rose from 80% to 85%. GDPval-AA v2 Elo climbed from 1615 to 1709 in xhigh and 1754 in max. The max tier gains ground through more turns and reasoning tokens: relative to xhigh, GDPval reasoning volume is approximately 62% higher and Tau3 is about 28% higher. Meta's self-reported DeepSWE v1.1 score is 75.4, compared to 55 for 1.2, 73 for GPT-5.6 Sol, and 74 for Claude Opus 5 in the reference table. These two measurement sets use different methodologies and cannot be directly summed or compared.

Social products are slated for updates soon, while a larger model remains in training. Version 1.3 will roll into Facebook, Instagram, and Meta AI applications within days, though no exact date has been provided. Alexander Wang described the pricing as "aggressive," noting that safety and alignment investments have been significantly increased, adding, "We haven't been forced to pause, but we need to make sure we don't hit the guardrails." A larger model codenamed Watermelon remains in development with no release timeline set. Zuckerberg has previously indicated that some generation of Spark weights might be open-sourced, but this version remains a closed, paid API.

Four releases in five months have moved Meta from a tier behind to a position where it now competes directly with Claude Fable, GPT-5.6, and Grok. Currently, the publicly available tier is xhigh, rather than the higher-performing max tier. Prices are held at 1.2 levels, but the increased token consumption from agentic tasks means real-world usage costs could rise. The next verification points will be when the max tier completes safety testing and opens, as well as the progress of actual replacement within social products.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10