On 15 September 2026, as Dreamforce opened — its big annual conference in San Francisco — Salesforce presented Koa, its first AI reasoning model, developed with NVIDIA. Salesforce publishes customer relationship management software, a CRM: the tool where salespeople track their prospects and where customer service teams handle requests. Koa is meant to power the AI agents the vendor sells under the name Agentforce, assistants able to act directly inside the CRM.
Since then, the announcement has been circulating on social networks with a spectacular story attached. Salesforce, “which is not an artificial intelligence player”, is said to have created an AI “stronger than GPT-6 Astra”, OpenAI’s latest model. OpenAI, Anthropic and Google are said to be “trembling”. And Koa is said to open a “second wave” of job cuts. We put each of those claims against the documents Salesforce published, starting with the technical report signed by its own researchers. The reality is less spectacular. It is also more interesting.
What Salesforce actually announced
Koa is not a model created from nothing. Salesforce started from Nemotron 3 Super, an open-weight model NVIDIA published on 11 March 2026. It has 120 billion parameters, but calls on only 12 billion to produce each word, which makes it cheap to run. Its weights are downloadable, it understands French and its licence allows commercial use.
Salesforce’s teams then retrained it through reinforcement learning: the model practises on simulated tasks and receives a reward when it solves them by calling the right tools. According to Salesforce’s statement, no customer data was used. The corpus is entirely synthetic. Scenarios covering more than 14 sectors, from manufacturing to healthcare by way of finance and travel, pair a profile type with tasks, then with the sequence of actions an agent has to chain to complete them. Jayesh Govindarajan, executive vice president for AI at Salesforce, told TechCrunch that those simulations staged everything from furious customers calling support to salespeople trying to close a deal.
The “27 years of CRM intelligence” Salesforce’s communications put forward therefore refers to know-how translated into scenarios, not to customer files. The technical report specifies that, for enterprise tasks, those scenarios are generated from agent descriptions written in Agent Script, Salesforce’s language for building Agentforce agents.
Koa runs in Salesforce’s infrastructure, which keeps the model weights. Unlike Nemotron, it is not downloadable: it is selected inside Agentforce, for a whole company or agent by agent, as one model among others. It already powers an internal Salesforce agent in Slack that helps employees find information, and it is entering a pilot phase with six customers: 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine and Xero. General availability is announced for winter 2026 in Salesforce’s American regions, with no date for the others, and will be followed by an open beta. No price has been disclosed.
The agreement goes beyond CRM. Salesforce and NVIDIA are extending it to Missionforce, the offering aimed at government and regulated sectors: retrained NVIDIA models are to be offered there to a few customers from October 2026, including inside networks cut off from the internet.
True, overstated or false: what is circulating about Koa
| What is circulating | What the sources say |
|---|---|
| "Salesforce is not an artificial intelligence player" | False. Its lab has already trained code, language, image captioning and speech recognition models, LeMagIT recalls, and the Koa report is signed by 25 researchers. What is new is an in-house reasoning model: until now, Salesforce handed that task to OpenAI or Anthropic. |
| "It is its own AI, not OpenAI's, Anthropic's or Google's" | True, with a clarification. Koa comes from none of those three labs, but it did not start from scratch: it is an NVIDIA model, Nemotron 3 Super, retrained by Salesforce. |
| "It was trained on all of Salesforce's data" | False. According to Salesforce, no customer data was used. Training rests on synthetic scenarios, including simulated angry customers and sales negotiations. |
| "It is stronger than GPT-6 Astra" | Not demonstrated, and contradicted by the published figures. No document compares Koa with GPT-6 Astra. Salesforce's technical report ranks it behind GPT-5.5 and Claude Opus 4.8. |
| "OpenAI, Anthropic and Google are trembling" | Overstated. Three weeks earlier, Salesforce widened its partnership with Anthropic, whose Claude model remains the default choice in several of its products. Koa is added to the models available, it does not replace them. |
| "This is the second wave that will destroy huge numbers of jobs" | Not established. Koa powers the same agents as before, at a cost presented as lower. The job cuts Salesforce attributes to AI in its customer support date from 2025, before Koa. |
”Stronger than GPT-6”: what Salesforce’s own figures actually say
Salesforce published two sets of figures. They do not tell the same story.
The first is commercial. The statement says that on CRM Bench, Salesforce’s in-house test suite (updating an opportunity, routing a request, scheduling a follow-up), Koa “matches or exceeds” leading models on CRM actions, “with three times fewer errors”. The product page adds that, against the “general-purpose models used by default today”, Koa is 11% more accurate at choosing the right action, 2.1 times more reliable at recalling a customer’s context and 15% better at holding the thread of a long conversation. None of those comparison models is named.
The second is scientific. The day before the announcement, on 14 September, 25 Salesforce researchers published Koa’s technical report. In it they set the model against three proprietary models and its own base model, on three tests: Tau2Bench, which simulates complete customer service conversations for an airline, a retailer and a telecoms operator; BFCL, which measures tool use; and CRM Bench.
| Model | Customer service (Tau2Bench, %) | Tools (BFCL, %) | CRM (CRM Bench, out of 1) |
|---|---|---|---|
| GPT-5.5 (OpenAI) | 84.0 | 67.6 | 0.90 |
| Claude Opus 4.8 (Anthropic) | 74.0 | 78.2 | 0.87 |
| Salesforce Koa | 69.4 | 66.6 | 0.86 |
| Nemotron 3 Super (NVIDIA, base model) | 68.6 | 64.7 | 0.84 |
| GPT-4.1 (OpenAI) | 54.5 | 54.0 | 0.81 |
Source: Koa technical report, Salesforce AI Research, 14 September 2026. Tau2Bench: share of customer service tasks completed, averaged across three domains weighted by task count. BFCL: overall accuracy in tool use. CRM Bench: average score on CRM tasks.
The verdict is clear. Koa improves on its base model, but barely: less than a point on end-to-end customer service, two points on tool use. It comfortably beats GPT-4.1, an OpenAI model released in spring 2025. And it stays behind GPT-5.5 and Claude Opus 4.8 on all three tests, nearly 15 points behind GPT-5.5 on customer service. Even on the criterion the product page leads with, choosing the right action, the report gives Koa an accuracy of 0.77 on CRM Bench, below that of the three proprietary models tested, which ranges from 0.82 to 0.85. The authors sum it up themselves: Koa exceeds “a strong proprietary baseline”, GPT-4.1, while remaining “below the strongest frontier models”.
The phrase “three times fewer errors” appears nowhere in that report. Silvio Savarese, chief scientist of AI research at Salesforce, explained on 16 September, according to LeMagIT, that the document deliberately did not present the internal test results, which could not yet be disclosed. According to LeMagIT, that internal test suite places Koa at GPT-5.5’s level for following instructions or summarising a text. On agent tasks, it is said to match Claude Sonnet 4.6, but to stay behind GPT-5.5, two variants of GPT-5.6 and Gemini 3.7 Flash — that is to say the very tasks Koa is meant to perform.
As for GPT-6 Astra, released on 3 September, twelve days before the announcement, it appears in no document Salesforce has published. The comparison going around has no basis. To measure what GPT-6 Astra is worth against its great rival, see our GPT-6 Astra versus Claude Fable 5.1 comparison. And to place Koa’s base model against the big general-purpose models, our model comparator lets you set Nemotron 3 Super against GPT-6 Astra or Claude, use by use.
The real story: almost as good, for far less
If Koa is not the best, why did Salesforce build it? Jayesh Govindarajan summed it up to TechCrunch: Salesforce had already created “many small specialised models”, but for reasoning the company had so far depended on the big labs. “Until now.” Before Koa, as soon as an agent had to run a long or multi-step task, the request went out to Claude or ChatGPT.
And the volumes are soaring. According to its quarterly results of 26 August 2026, Salesforce counted 3.2 billion “agentic work units” across Agentforce and Slack in the second quarter, almost double the previous quarter, and Agentforce passes $1.5 billion in annual recurring revenue, a measure that now includes Slackbot and other AI offerings. At that scale, every task handed to an outside model is paid for per token, the unit of text a model reads or writes.
That is where Koa scores. According to LeMagIT, the internal test suite presents it as the cheapest to run of the models compared, and Jayesh Govindarajan says its cost-to-accuracy ratio is “markedly better”. At NVIDIA, Kari Ann Briski talks about “tokenomics”, the economics of the token. The training itself was frugal: 32 B200 graphics processors, according to LeMagIT, where Thomson Reuters used 368 to train its own model, Thomson-1.0-Large.
Second argument, control. Salesforce trains and runs Koa in-house, and no customer data crosses its walls, neither during training nor during use. For a bank, a hospital or a government body, that is often more decisive than two points of score.
Third argument, provenance. Why Nemotron, and not a Chinese model? Jayesh Govindarajan is explicit: before Nemotron there was, in his view, no American sovereign base model that was simultaneously available, at the state of the art and of clear data provenance. “We have no idea what Qwen is trained on,” he added, about Alibaba’s open model. The jab points at a reality: Chinese open models weigh heavily in global usage, as we showed in explaining why China gives its best AI away free. NVIDIA is positioning itself as the American alternative.
OpenAI and Anthropic are not trembling yet: Salesforce is playing both sides
On 26 August 2026, three weeks before Koa, Salesforce and Anthropic announced Claudeforce. The agreement works both ways. Salesforce moves into Claude as an extension with 37 ready-made sales skills, and Claude serves as a reasoning model inside Agentforce. According to the statement, Claude is above all the default model for Slack AI, Slackbot and Agentforce Coworker, and Claude Code is used across Salesforce engineering.
Koa therefore replaces nothing. It joins the list of models customers can choose inside Agentforce. TechCrunch makes the point: Salesforce is “not really abandoning” Anthropic or OpenAI. The logic is one of routing: repetitive, high-volume tasks to a cheap in-house model, difficult cases to the frontier models.
The signal to the labs is real nonetheless. Salesforce is not the first to do this arithmetic. In May 2025, ServiceNow, another enterprise software vendor, was already presenting with NVIDIA the Apriel Nemotron 15B, a compact reasoning model for IT, human resources and customer service, pitched as faster and cheaper to run than giant general-purpose models. At Dreamforce, Jensen Huang, NVIDIA’s chief executive, said every company and every country would become an AI player, and that the share of open models had gone from 30% in early 2025 to “some 70%” today, without specifying what that figure measures.
What the labs risk losing is the steadiest part of corporate spending: the millions of small repetitive tasks. For now, they keep what Koa cannot do as well. The surest winner in the operation is NVIDIA, which supplies the base model, the training tools and the processors.
Jobs: a “second wave”? What is established, and what is not
Koa targets precise tasks: generating leads, qualifying sales opportunities, resolving customer service requests. These are pieces of the daily work of salespeople, sales assistants and customer advisers. Should we see a new wave of job cuts in it? Three facts counsel caution.
The wave has already happened at Salesforce, without Koa. At the end of August 2025, Marc Benioff, Salesforce’s chief executive, said he had cut customer support headcount from 9,000 to around 5,000 people, “because I need less heads”, CNBC reports. The company explained that the number of requests to handle was falling and that it no longer needed to actively replace its support engineers. At the time, its agents’ reasoning went through OpenAI or Anthropic models.
The 2026 cuts are not officially attributed to AI. In August, Salesforce notified authorities of 133 further job cuts in Washington State and San Francisco, effective 5 October, its third plan of the year. According to The Next Web, the mandatory filing lodged in Washington State does not mention AI.
Agents are not reliable on their own. On Tau2Bench, which simulates complete customer service conversations, Koa succeeds on an average of 69% of tasks, and the best model tested, GPT-5.5, on 84%. In other words, the best agent still fails on roughly one task in six, and Koa on nearly one in three. At that level, you need humans to take over. Our test of automated AI prospecting reached a similar conclusion: AI speeds up research, writing and follow-ups, but it guarantees neither meetings nor qualified leads without human follow-through.
What Koa really changes is the price of an automated task. If a follow-up or a request triage costs less to hand to a machine, companies will automate more of them. That can remove posts. It can also get tasks done that nobody had time for. The available data mainly shows an effect on early-career posts, as we set out in our review of the studies on jobs at risk from AI, and some of the companies that replaced staff with AI have already reversed course: that is the reversal behind layoffs “because of AI”. In France, finally, technological change is written into law among the grounds for economic dismissal, the opposite of what several Chinese courts have ruled.
What happens over the coming months
- September and October 2026: the pilot phase. Six customers are testing Koa. Salesforce says it is already open, NVIDIA places it in October. Missionforce opens retrained NVIDIA models to a few government bodies and regulated organisations at the same time.
- Winter 2026: general availability, in Salesforce’s American regions, then an open beta.
- Europe: nothing announced. European Salesforce customers will have to wait, and check where the model will be run.
- The missing figures: Koa’s price, the internal test results, and above all a first independent evaluation. As long as that is missing, Koa’s performance remains what its vendor says it is.
Longer term: three trends to watch
The end of the single model. Software vendors will increasingly assemble several models: an in-house model for volume, a frontier model for hard cases. Koa’s report says it in black and white: open models have closed much of their gap on public tests, and access to their weights lets companies specialise them to better control deployment, governance and cost.
Pressure on lab pricing. It will bear first on repetitive tasks, where a “good enough” model suffices, without challenging the lead of OpenAI, Anthropic or Google on complex tasks.
A sovereignty question. Salesforce posed it for the United States by ruling out a Chinese model. Europe is already asking it, as Mistral’s choice to host a Chinese model showed.
Le Recul’s reading
Koa is neither a superintelligence nor a scientific revolution. It is an industrial decision: own a slightly less good model, but cheaper and entirely controlled, rather than rent the best model on the market for every task. The viral story inverts the order of things. Koa does not threaten OpenAI because it is stronger, but because it is good enough for a large part of everyday work.
That is also what makes it a real employment story, with one nuance: it is not model capability that decides job cuts, it is cost. And the two figures that would let anyone judge — Koa’s price and its internal results — are not published.
What to take away
- Koa is Salesforce’s first reasoning model, built on NVIDIA’s Nemotron 3 Super and retrained without customer data.
- It has never been compared with GPT-6 Astra. Salesforce’s report ranks it behind GPT-5.5 and Claude Opus 4.8, ahead of GPT-4.1.
- Its interest lies elsewhere: a cost presented as far lower and total control of the data, for the repetitive tasks of sales and customer service.
- Salesforce is not leaving Anthropic: Claudeforce, announced three weeks earlier, remains at the heart of its products.
- For jobs, Koa extends a trend already under way, with customer support going from 9,000 to around 5,000 people in 2025, with no evidence of a “second wave”.
- General availability in winter 2026, in the United States only. Nothing is announced for Europe.