What is Muse Spark 1.1, and should you switch to it?

Muse Spark 1.1 Changes Everything... — watch on YouTube
9 min

In 30 seconds

  • Muse Spark 1.1 is built for agent work. It delegates to parallel copies of itself and clicks through real browsers on desktop and mobile.
  • At $1.25 in and $4.25 out per million tokens it costs a quarter of Claude Opus 4.8 on input and a sixth on output, which is the line item that grows when an agent runs all day.
  • Every published score comes from Meta's own launch table, there is no independent score for 1.1 yet, and Meta was caught inflating Llama 4's benchmarks in 2025.

Muse Spark 1.1 is Meta’s newest AI model, built to do work rather than chat. It runs teams of copies of itself, drives a browser the way a person does, and costs a quarter of Claude Opus 4.8 on input tokens and a sixth on output.

What is Muse Spark 1.1?

The name decodes the same way Claude’s does. Muse is the family, Spark is the model line inside it, and 1.1 is the update to 1.0, which shipped in April. It is fully multimodal with a one million token context window, matching Opus 4.8 and OpenAI’s GPT-5.5. Meta has not published the parameter count or the architecture. You can use it free in the Meta AI app today, and over the next few weeks it is expected to replace the older models running the AI inside WhatsApp, Instagram and Facebook. Meta also launched the Meta model API, the first time outside developers can pay Meta to build on one of its models.

Two things stand out in how it works.

First, it delegates to itself. A main copy plans the big job and hands pieces to parallel copies called sub-agents, the way a lead engineer splits the day’s calls across the vans and keeps the whole schedule in his head.

Second, it operates a computer the way you do, clicking interfaces and typing into browsers on desktop and mobile, and it decides for itself whether a script or the screen is faster.

Meta showed one demo worth watching. You film a short video of something you want to sell, an old chair or a bike. The model watches the video, pulls the best frames, works out what the product is and what it is worth, then opens your browser and builds the Facebook Marketplace listing, photos and title and description and price included. Nothing about that sequence is specific to Facebook Marketplace.

Why did Llama 4 fail, and what Meta spent to recover

Llama, Meta’s open model line, ran for years under Yann LeCun, hired in 2013 to build the FAIR lab on the condition that whatever he built stayed open and free. Llama 4 was meant to put Meta in front. It flopped in April 2025 because the team was caught inflating its own benchmark numbers, submitting one model to public leaderboards and shipping a different one to developers. On the independent Intelligence Index, Llama 4’s best version scored 18 while the leaders sat in the high 50s.

Meta then paid $14.3 billion for a 49% non-voting stake in Scale AI, mostly to get its 28-year-old founder Alexandr Wang as chief AI officer, and raided OpenAI, Anthropic and Google DeepMind at $100 million to $300 million a head. Some researchers ignored Zuckerberg’s DMs because they assumed they were fake. LeCun left, over 600 people lost their jobs, and in April 2026 the new lab shipped Muse Spark 1.0 at a score of 52 on the index where Llama 4 scored 18.

How much does Muse Spark 1.1 cost?

Model Input, per million tokens Output, per million tokens
Muse Spark 1.1 $1.25 $4.25
Claude Opus 4.8 $5 $25
Fable 5 $10 $50

Input is what an agent reads, output is what it writes, and both are billed per million tokens. Reading is where the volume sits. An intake agent that goes through every inbound job request, every photo caption and every previous note on the customer burns input tokens all day and writes back three lines. At $1.25 against $5 that day of reading costs a quarter as much, and on output the gap is wider, $4.25 against $25.

That is the whole cost story, and it only shows up on a usage bill rather than a subscription. If your automation runs on a per-seat plan, the table above changes nothing for you this month.

Muse Spark 1.1 vs Opus 4.8 and GPT-5.5: the benchmarks

Model Humanity’s Last Exam, with tools SWE-Bench Pro
Muse Spark 1.1 62.1 61.5
Claude Opus 4.8 57.9 69.2
GPT-5.5 not given behind 61.5
Fable 5 not given 80.3
GLM 5.2 not given 62.1

Humanity’s Last Exam with tools is a brutal reasoning test where the model gets search and code, and Muse Spark 1.1 beats Opus 4.8 on it. It also takes the top spot on the two big tool-use benchmarks, MCP Atlas and Job Bench. Hard coding is weaker. Its 61.5 on SWE-Bench Pro ties GLM 5.2, the open Chinese model you can download and run yourself for nothing, and on terminal work it comes third behind GPT-5.5 and Opus. Best at orchestrating tools and running errands, a step behind at writing hard code.

Every number above comes from Meta’s own launch table, and Meta chose which modes to show its rivals in. This is the company that admitted juicing Llama 4’s benchmarks. The eval report is public and independent testers will have their own figures within days, but as of today there is no independent score for 1.1. The full set Meta published, including the safety numbers, is in the technical breakdown of Muse Spark 1.1.

Should you switch to Muse Spark 1.1?

Not this quarter. Watch the computer use rather than the price.

Take a warranty claim on a compressor that failed inside its cover. Someone in your office logs into the manufacturer’s portal, types the serial number off a photo the engineer sent from the roof, picks a fault code from a dropdown, uploads the invoice, submits, then pastes the claim number back into the job so the customer can be told. That portal has no API and never will, which is why every automation quote you have had either skips that job or prices a person to sit and do it. A model that clicks interfaces and reads photos runs the whole sequence, and it is the Marketplace demo with a different website in front of it. Supplier reordering, utility interconnection paperwork and the insurer’s claims site all look the same from the model’s side.

So the useful work this week is a list. Write down every job that stalls because the software involved has no API and someone has to type into a screen. That list is what an agent with computer use is worth to you, and it does not change if the benchmark table turns out to be generous. Then ask whoever built your automation what it spent on tokens last month, and what that same month costs at a quarter of the input rate.

Two things are true about Meta regardless. Wang says a model trained on ten times the compute behind Muse Spark 1.0 is already matching GPT-5.5, with an Opus-level coding model coming soon. And Meta ships this model as the default brain inside apps that more than 3 billion people already use, so your customers will be talking to it inside WhatsApp whether or not you ever call the API. Write the list, get last month’s token bill, and check the independent scores when they land in a few days.

Full transcript

Expand

Meta just dropped their newest AI model, Musepark, and the internet is going crazy. And the past 10 days alone have possibly been the craziest 10 days in AI history. Anthropics Fable 5 got on back. XAI turned into SpaceX AI and they dropped a new Grock and OpenAI released GPT 5.6. We've had four Frontier models hit the market in the last month alone.

But in this video, we're going to talk about Muse Spark and Meta because I believe they're going to be the dark horse of the AI race. From the stats alone, this model is looking very impressive. So, in this video, we'll break down exactly what Muspark 1.1 is, how Meta has just become a new powerhouse contender in the AI race, and how they're actually stacking up head-to-head with the big boys. So, let's dive right in. So, what actually is Muse Spark 1.1? First, let's decode the name because like AI companies do, Meta has also made the name of their model very confusing. Now, Muse is the family like Claude is for anthropic and Spark is the model line inside that family like Opus is for Claude and 1.1 like you would have guessed it is the newest version to 1.0 which came back in April. Now, Meta's own framing is that this model has been created with Agentic Tasks in mind. You can use it right now for free inside the Meta AI app. And over the next few weeks, it's expected to replace the older models powering the AI inside WhatsApp, Instagram, and Facebook. Now, the real headline here is actually in how Meta is selling it.

Alongside Muse Park 1.1, Meta has also launched Meta model API. And this is the first time that outside developers can actually pay Meta to build on top of one of its models, which is crazy because it's a big shift in what Meta has been doing over the last decade or so when it comes to AI and their original Llama series. If you rewind to 2013, Meta, which was just Facebook back then, set up one of the first serious AI labs in the world called Fair. And to run it, they hire Yan Lun. I apologize if I'm butchering that name. I don't speak French very well. But what you must know is that he is literally considered to be one of the three men known as the godfathers of modern AI. A little under a decade later when the chat GBT wave hits, Lacon agrees to build Meta's own language model on one condition so long as it stays open and free. That model became Llama and Llama blew up, becoming pretty much the most widely used open weights AI model in history. Then came April 2025 and Llama 4. Llama 4 was supposed to be the model that put Meta at the front of the AI race, but instead it flopped. And it flopped extremely hard because Meta's team got caught juicing up their own benchmark numbers.

And the version they submitted to public leaderboards wasn't even the same model they shipped to developers. Now, to put the actual gap in one number, there's an independent scale called the Intelligence Index that scores every major model. On that index, Llama For's best version scored an 18 while the leaders at the time were sitting in the high50s. Now, this was obviously a huge liver shot to our main man Zuckerberg.

So, our guy decided to go and plan the most expensive comeback in AI history. His first move was to pay $14.3 billion for a 49% non- voting stake in a company called Scale AI in June of 2025. Now, the real prize here wasn't actually just their company. It was their 28-year-old founder, Alexander Wong, who Zuckerberg made Meta's chief AI officer. And then came the talent rate. Meta went after the top researchers at OpenAI, Enthropic, and Google DeepMind with pay packages that reportedly ran from 100 to $300 million each. Zuckerberg himself was reportedly running recruiting personally, and funny enough, some AI researchers just thought his DMs were fake and just ended up ignoring him.

Now, all this talent got rolled into a new division called Meta Super Intelligence Labs with Wong running the frontier models and former GitHub CEO Nat Freiedman running products. But this comeback also left over 600 people out of jobs. And Lun himself actually also left the company after being sidelined by Wong. Less than a year later after their Llama 4 crash out in April of 2026, that lap shipped his first model Muspark 1.0. And the model was good. On that same intelligence index where Llama scored 18, Musepark 1.0 scored 52. Which brings us to today and what Muspark 1.1 actually brings to the table. Let's start with the raw specs. Muspark 1.1 has a 1 million token context window which is on par with Clausopus 4.8 and OpenAI's GBT 5.5. It's fully multimodal, but Meta has still not disclosed the parameter count or the architecture. But the headline feature is how it works as an agent. And two things stand out.

First, it's built to run teams of itself. As a main agent, it can plan a big task and delegate pieces to parallel copies of itself called sub aents all working at once. Second, it's trained to use a computer like you do. It can click on interfaces, type into browsers, and work on desktop and mobile. And it actually has a level of judgment because it can decide whether it's faster to write a script versus just clicking through the screen. And the demo that actually shows what that means in real life is actually very impressive. You take out your phone and you shoot a quick video of something you want to sell, say an old chair or a bike or whatever. Muse watches the video, pulls out the best frames, figures out what the product is and what it's worth, then opens your browser and builds the Facebook Marketplace listing for you.

We're talking photos, title, description, price, all of that just from you filming a video. And then there's the cost of using the model. Muse 1.1 costs $1.25 25 per million input tokens and $4.25 per million output. For comparison, Opus 4.8 is $5 in and 25 out. Fable 5 is $10 in and50 out. So Meta is undercutting the entire American Frontier market by quite a margin. We've talked about a couple features so far, but how good is Museark 1.1? Actually, here's the pattern. On humanity's last exam with tools, which is a brutal frontier reasoning test where the model gets to use search and code, it scored 62.1, beating clonopus 4.8's 57.9. It also took the top spot on the major tool use benchmark like MCP Atlas and Job Bench. Now, where it's not quite as good is raw coding. On the SweetBench Pro, it scored 61.5. Now, that edges out GBD 5.5, but it's well behind Opus 4.8 69.2 2 and it is nowhere near Fable 5's 80.3. Now GLM 5.2 which is the open Chinese model we've covered in the past scores 62.1 on the Sweepbench Pro. So statistically that's a tie against the model you can download and run for free. And on terminal work it's a clear third behind both GBD5.5 and Opus. So it's bestin-class at orchestrating tools and running errands but a little bit behind on writing hard code. Now one quick disclaimer. Every single number I just gave you from Espark comes from Meta's own launch table. And on that table, Meta showed which modes to show its rivals in. This is the company that by its own admission juiced up benchmark numbers for Llama 4.

As of the release of this video, the Eval report is public and independent testers will verify them in a couple of days. But as of today, there is no independent score for 1.1 yet. So, what does this mean for Meta? There are three things you should know about for the foreseeable future. First is how these new gen models actually progress.

Alexander Wong says Meta is already training a new model using 10 times the compute power behind Mu Spark 1.0 and he claims it already matches GPD 5.5 with an Opus level coding model coming very soon. There's also Muse Video which is already in development and Muse image which released a couple days ago. Now the second thing is the money coming in.

Meta now has a paid API and consumer subscriptions. Not to mention the quiet giant underneath all of this. Better AI capabilities means better ad targeting and more engagement across all platforms that Meta hosts. Primarily Facebook and Instagram where Meta makes its billions every single year. Number three is automatic adoption. Meta ships this model as the default brain inside the apps that more than 3 billion people already use. This is the Gemini play, except the average person today probably even spends more time on meta apps than they do on Google and YouTube. So, I just don't see how Meta is not going to be a serious player in the AR race for the years to come. And I think the gap is only going to get more narrow as the months go on. So, that's MSAR 1.1 for you. If you enjoyed this video, you can subscribe to the channel. If you want to get in touch with me personally, you can get in touch with me on LinkedIn. Link is down in the description below. And if you run a business doing anywhere from $2 to $10 million in annual revenue, you can book an AI opportunity audit with myself and my team. That is the first link down in the description. I'll see you in the next

Next in The AI BriefingHow I Plan to Make +200 Videos / Month With AI agents
Ready to put AI employees to work?
Book a call
256-bit SSL Secured
© 2026 Nairon, Inc. All rights reserved.
PrivacyTerms & ConditionsCookie PolicyAcceptable Use