On 10 August 2026, Meta released Muse Glimmer: a thirty-billion-parameter artificial intelligence, under the Apache 2.0 licence, designed to run on a personal computer. It is its first open-weights model since Llama 4, sixteen months earlier.

Four days before that, on 6 August, OpenAI moved ChatGPT to unlimited text for every free account.

Two opposite ways of paying nothing, four days apart. And a question millions of people ask without ever getting a quantified answer: has the AI running on my machine caught up with the one in the cloud?

The answer is no. Here is by how much, what it really costs, and the one reason people do it anyway.


Results: the verdict in brief

On intelligence, free ChatGPT wins comfortably. Muse Glimmer sits at the level of OpenAI’s free model at its lowest setting, and nothing beyond.

On hardware, the entry ticket is higher than advertised. The model “fits on a personal computer”, but on an ordinary Mac it produces three words per second. For it to become comfortable, you need a graphics card costing the equivalent of four and a half years of a ChatGPT subscription.

On data, the match is one-sided in the other direction. Nothing leaves your machine. Facing it, on the free tier, your conversations are used by default to train the model.

And there is one area where local wins for an unexpected reason: agentic work. Agent mode is paid at OpenAI, absent from the free tier — and it is exactly what Meta designed its model for, and it is given away.

This is a match between two things that do not compare term for term — which is precisely why it had to be run.


Where these numbers come from

Method before figures. Every data point can be replayed line by line:

  • Intelligence scores come from the Artificial Analysis index, read setting by setting for both models, so that the comparison covers equivalent quantities rather than two incomparable averages.
  • Memory requirements, throughput and failure modes are hardware measurements, attributed to the card concerned, distinguishing what is measured from what is estimated. AMD figures come from the manufacturer, RTX and Apple figures from independent reports.
  • The content of ChatGPT’s free tier is established from OpenAI’s documentation and a specialist press inventory.
  • Break-even calculations are redone from observed prices, with the assumptions written into the text.

Two points are not measurable as things stand, and we flag them rather than fill them: quality in languages other than English, for which Meta publishes no evaluation, and the exact caps on images, files and voice on the free tier, which OpenAI does not publish.


The two contenders, side by side

Muse Glimmer 30B Free ChatGPT (GPT-5.6 Luna)
Vendor and date Meta — 10 August 2026 OpenAI — free tier switched on 6 August 2026
Where it runs On your machine On OpenAI's servers
Size 30bn parameters, dense transformer not published
Context 131,072 tokens 1,000,000 tokens (8 times more)
Inputs text, image (built-in perception encoder) text, image
Knowledge cutoff 4 January 2026 not published (web search included)
Licence Apache 2.0, commercial use free a service, under terms of use
Cost €0 in licence, but a machine €0, but your data

One useful clarification before going further: Muse Glimmer is distilled from Muse Spark, Meta’s proprietary model. This is not a model cobbled together at the margins, it is a condensed version of the in-house flagship.


Intelligence: the gap, measured on a common index

This is the heart of the match, and the comparison only makes sense if you go down to the setting level. A reasoning model does not have one intelligence, it has as many as it has effort levels.

Model and setting Intelligence index
Free ChatGPT — low setting 34
Muse Glimmer (on your machine) 35
Free ChatGPT — medium setting 39
Free ChatGPT — high setting 47
Free ChatGPT — very high setting 50
Free ChatGPT — maximum 52

The result fits in one line: a free artificial intelligence running on your own computer today reaches the weakest setting of free ChatGPT, and nothing above it.

That is not a humiliation — it is remarkable for thirty billion parameters fitting into seventeen gigabytes. Muse Glimmer is in fact five index points ahead of Gemma 4 31B, and Artificial Analysis notes that it matches a model with thirty-three times more total parameters. Against other open models of its size, it leads. Against OpenAI’s free tier, it plateaus. Our ranking of open models is where that comparison is fair.

On agentic evaluations it holds up far better: MCP Atlas 75.5; DeepSearch QA 74.6; GAIA2 43.3; SWE-Bench Pro 51.2; AIME 2026 94.7; IFBench 77.0. That is where Meta concentrated its effort — agents that hold a long task, call tools and recover from an error. Qwen is ahead of it on OSWorld-Verified, 75.6 against 65.9, and on most multimodal evaluations.


What machine you need: the real entry ticket

“Runs on your computer” is true. What the phrase does not say is how fast.

Hardware Memory Observed price Throughput
Entry-level M3 / M4 Mac 24 GB unified already owned 4.3 tokens/s (≈ 3 words/s)
AMD Ryzen AI Max+ 395 unified whole machine up to 24 tokens/s
Apple M4 Max unified — 23.7 → 37.8 tokens/s (×1.5)
RTX 3090 24 GB ≈ $1,248 second-hand 65-70 tokens/s (estimated)
RTX 4090 24 GB ≈ $2,268 second-hand 75 tokens/s (measured)
RTX 5090 32 GB $4,399 and up 74.9 → 233.4 tokens/s (×3.1)

Remember the first and the last line: 4.3 tokens per second against 233.4. A factor of fifty-four, for the same model.

On an ordinary Mac, Muse Glimmer works — you watch the text appear at roughly the speed of attentive reading. That is usable for a summary, painful for a conversation, out of the question for an agent chaining dozens of calls.

One point deserves crediting to Meta: memory efficiency is remarkable. An independent measurement on an RTX 4090 gets 130,000 tokens of context in 19.3 gigabytes total, with no attention-cache compression. The attention scheme chosen is unusually frugal, and that is real engineering.


The DFlash trap: why the three-times speed-up is not for you

Meta highlights a spectacular figure: from 74.9 to 233.4 tokens per second, a 3.1× multiplier. The technique is called DFlash, a speculative decoding scheme: a five-layer, 5.11 gigabyte companion model proposes blocks of sixteen tokens in a single pass, which the main model verifies in parallel.

Do the arithmetic.

17.3 gigabytes for the model, plus 5.11 for the companion, equals 22.4 gigabytes — before the attention cache and before the vision encoder. On a 24-gigabyte card of which roughly 23.5 are actually usable, that leaves about one gigabyte. Which is a few thousand tokens of context, and no images.

The three-times speed-up is therefore only reachable from 32 gigabytes, meaning an RTX 5090 at $4,399. On 24-gigabyte cards — the ones Meta names as its target — the companion, the cache and the encoder fight over about six gigabytes that do not exist.

Apple unified-memory machines partly escape the problem and gain 1.5 times instead of 3.1. Less spectacular, but real.


Three ways to get the install wrong

These are the three recurring mistakes, and two of them fail silently — the worst possible way.

The 16-gigabyte card. The four-bit file weighs 17.3 gigabytes. It does not fit. Ollama then loads what fits and quietly spills the rest onto the CPU: on a dense thirty-billion-parameter model, speed falls to single-digit tokens per second. Nothing signals the error, the model simply answers very badly.

Speculative decoding on 24 gigabytes. Same mechanism: enabling the option causes a spill onto the CPU. The check is simple — the ollama ps command must show 100% GPU in the processor column. Any other value signals the spill.

The Ollama version. 0.32.7 failed outright on NVIDIA and AMD hardware. 0.32.8 is required outside Apple machines.

As for choosing a quantisation, it comes down to this: the official four-bit version weighs 17 gigabytes and targets 24 GB cards; the community five-bit variant rises to 20.1 and eats into context space accordingly; the eight-bit one reaches 29.6 and demands a 32 GB card. Meta’s official dynamic quantisation, reserved for 32 GB, claims 0.2% accuracy difference from full precision.


What free ChatGPT actually gives you since 6 August

Facing it, the offer changed in kind on 6 August 2026 — and it is less generous than the word “unlimited” suggests.

What is unlimited: text, and text alone. Free accounts send as many text messages as they want, but are locked to a single model, GPT-5.6 Luna, the smallest of the GPT-5.6 family.

What is not locked, though: the reasoning level. The 6 August announcement also gives free accounts a “Think” button, which grants the model more reasoning time on harder questions, at no extra cost and “subject to anti-abuse guardrails”. OpenAI does not say which reasoning level it corresponds to — a silence that weighs, given the gap measured above between the low setting (34) and maximum (52). We dug into that point in our comparison of DeepSeek V4.1 Flash against GPT-5.6 Luna, where the same model is measured setting by setting.

What remains capped:

  • image generation, with a daily limit OpenAI does not publish and adjusts with server load;
  • file uploads, with a library limited to 500 megabytes;
  • voice, available but switched to a lighter model;
  • deep research and agent mode, simply reserved for paid plans.

Web search, on the other hand, is included — and that is a structural advantage Muse Glimmer does not have: its knowledge stops on 4 January 2026, permanently, unless you wire up a search tool yourself.


The reversal: agent mode is paid at OpenAI, free at Meta

Here is the point that inverts part of the verdict, and that no comparison picks up.

On ChatGPT’s free tier, agent mode and deep research are simply absent: they are paid features. A free user has a conversational assistant, not an executor that chains tasks.

And that is exactly what Muse Glimmer was designed for. Meta did not optimise it for conversation but for permanent local agents: tool calling, task memory, failure recovery, long sessions. And its results on those evaluations are not anecdotal — MCP Atlas 75.5, the tool-dialogue evaluation, DeepSearch QA 74.6, GAIA2 43.3.

The practical consequence is clear. On agentic ground alone, the relationship inverts: the free local model does what the free online one does not — not because it is better, but because the other one charges for it.

To which add the total absence of any quota. An agent running six hours straight on your machine consumes nothing but electricity. The same agent on an online programming interface is billed per token, and runs into caps. If that is your use, our comparisons of OpenClaw against Hermes Agent and of Codex against Cowork cover the agents you would plug this model into.

That is the nuance that stops you concluding too fast: free ChatGPT wins the assistant match, Muse Glimmer wins the executor match.


How to install it, concretely

A comparison that recommends a local model without saying how to get it does half a job. The common route takes a few minutes.

The model is distributed officially on Ollama, under the name muse-glimmer, with its quantised variants. It is also available on LM Studio, which offers a graphical interface for anyone who would rather avoid the terminal, and on Hugging Face for whoever wants the raw weights.

Three checks before launching anything:

  1. Available memory. Count 18 gigabytes minimum of genuinely free RAM or VRAM — not the machine’s total memory. On a unified-memory Mac, the system already takes a share.
  2. The Ollama version. 0.32.8 or newer outside Apple machines, otherwise loading fails on NVIDIA as on AMD.
  3. The quantisation. Take the official four-bit version if you have 24 gigabytes; the eight-bit version, more faithful, requires 32 gigabytes.

After the first launch, one check is worth doing: ollama ps must show 100% GPU in the processor column. Any other value means part of the model is running on the central processor, and that you are measuring a speed that has nothing to do with the one this article is about.


Your data: the one area where the match is lopsided

Up to here, ChatGPT wins almost everywhere. Here, the gap is absolute, and it runs the other way.

Muse Glimmer sends nothing. Once the weights are downloaded, everything runs on your machine: no request leaves, no connection is needed, no quota exists, and nobody — not Meta, not an internet provider, not a third party — can read what you write. That is not a privacy policy, it is a physical property of the installation.

On ChatGPT’s free tier, training on your conversations is on by default. OpenAI’s documentation provides that exchanges are used to improve the models; the company also collects IP address, location and session data. Opting out exists but it is manual, in the settings, under Data Controls. Deleted conversations can remain up to thirty days on OpenAI’s systems. And the free tier carries no data processing agreement, which disqualifies it for any professional use touching client data.

The European context is not neutral: in January 2026, the Italian data protection authority fined OpenAI 15 million euros for unlawful processing of personal data.

That is the real dividing line between the two: with free ChatGPT you pay in data; with Muse Glimmer you pay in hardware and in slowness. The question is never only “is this model good”, but “where does what I send it end up”.


What Muse Glimmer cannot do

Independent reviews converge, and Meta does not contradict them.

Professional writing, open questions and dialogue lag behind the large proprietary models. That is not what it was designed for: Meta optimised it for permanent local agents — tool calling, task memory, failure recovery.

The guardrails are less robust, with a higher violation rate in adversarial testing. A model you install yourself is also a model nobody is watching.

Multimodal capabilities are described as immature against models specialised in vision or document reading — even though it embeds a perception encoder.

And it is explicitly not recommended as an unsupervised authority for financial, legal, medical, safety or production decisions. For nuanced writing or regulated use, expert tuning is required.


Languages: what Meta says, and what it does not

Meta states that it trained Muse Glimmer on data covering more than a hundred languages. The same card specifies, in the same paragraph, that the model has not been evaluated across all of those languages and that performance may degrade outside the solidly supported subset.

Meta publishes neither the list of that subset, nor a single measurement in any language other than English.

That is careful wording, and it should be read for what it is: a hundred languages in training does not mean a hundred languages guaranteed. If you work in English, you are almost certainly inside the supported subset. If you do not, you have no way of knowing — no public comparative evaluation exists to date, which makes it impossible to settle other than by use, and an impression of use is not a measurement.


What it really costs, on both sides

The most useful calculation is the break-even point, and it gives a counter-intuitive result.

Against free ChatGPT, there is no break-even point. You will never pay off a hardware purchase against a service that costs zero. The financial comparison only makes sense against a subscription.

Against ChatGPT Plus, at around €23 a month:

Card Observed price Equivalent in subscription
Second-hand RTX 3090 ≈ $1,248 ≈ 54 months, or four and a half years
Second-hand RTX 4090 ≈ $2,268 ≈ 99 months, or more than 8 years
New RTX 5090 $4,399 and up ≈ 191 months, or nearly 16 years

And even then: those durations ignore electricity, the rest of the machine, and the fact that in four and a half years the landscape will have changed several times over.

The conclusion is therefore clear: you do not buy a graphics card to save money on ChatGPT. You buy it because you want nothing to leave your home, or because you run agents continuously, or because you develop and need weights you control. Three good reasons — none of which is price.


What our own ranking says

Muse Glimmer belongs to one precise category of our AI ranking: models you can run on a local machine. That is where it should be judged, not against frontier models served from data centres.

One honest clarification about that category: in practice it lists models that fit on a personal machine, with a size limit. A three-hundred-billion-parameter model does not appear there, whatever its licence. That is exactly Muse Glimmer’s ground — and the only one where the comparison is fair to it.

To place the other camp, our comparison of the two frontier models released a day apart in August shows the level at which the online race is being run. The gap with local remains considerable.


The story behind it: Meta closed, then reopened, in four months

This detail changes how you read the announcement, and nobody connects it.

In April 2026, Meta released Muse Spark, its first closed proprietary model — the end of the open-weights Llama era, after years in which the company had made itself the standard-bearer of open AI. Four months later, in August, it publishes Muse Glimmer under Apache 2.0, the most permissive licence in the industry.

Meta therefore closed, then reopened, in a single quarter. And Muse Glimmer is distilled from Muse Spark: it is the closed model that feeds the open one.

The simplest reading is also the most likely: open and closed are no longer two camps but two tiers of the same product line — the high end stays closed and paid, its condensed version is given away to hold the ground. We saw the same strategy from the opposite direction in our comparison of Kimi K3 against Claude Fable 5, where the open model is the cheap one and the closed one is the expensive one.


Which to choose for your situation

You just want good AI, free, with no effort. Free ChatGPT, without hesitation. Better, faster, kept current by web search, no installation. Take three minutes to disable training in Data Controls — that is the one gesture that matters.

You work on confidential documents. Muse Glimmer, and it is the only honest answer of the two. ChatGPT’s free tier has no data processing agreement and trains by default. No setting replaces the fact that nothing leaves the machine.

You run agents continuously. Muse Glimmer, provided you have 32 gigabytes of video memory. That is what it was designed for, its agentic scores confirm it, and the absence of any quota becomes decisive when an agent runs for hours.

You have an ordinary Mac and you are curious. Try it, but know you will get 4.3 tokens per second — about three words per second. That is an interesting discovery, not a working tool.

You are hesitating about buying a graphics card for this. Do not do it to save money: the cheapest suitable one is equivalent to four and a half years of subscription. Do it if privacy, offline autonomy or development justify it. You can compare the models themselves in our free comparator.


What to remember

No, local AI has not caught up with the cloud. Muse Glimmer reaches the weakest setting of free ChatGPT, has eight times less context, and requires a machine most people do not have. On an ordinary computer, it writes three words per second.

But the real lesson of this match is elsewhere, and it lies in a symmetry: both of these free offers are paid for, simply not in the same currency. One is paid for in data — training by default, thirty days of retention, no data processing agreement. The other is paid for in hardware and in patience.

The real progress of summer 2026 is therefore not that local has caught up with the cloud. It is that local became, for the first time, a defensible choice for reasons that are not performance.