Stories

Perplexity's agent runs on a local PC, and Nvidia pulls Claude off sensitive work

Google's new voice models switch between 97 languages mid-call, and Salesforce's support team is down from 9,000 people to about 5,000.

By , Senior AI ConsultantEdition of

6stories
5minute read
In This Edition

Perplexity's Computer agent now runs on the PC in front of the person using it, on one condition: the machine needs an NVIDIA RTX card with at least 24 gigabytes of video memory. It reads the files on that computer, plans a task of several steps and works through it, and what it reads is never sent to Perplexity's servers. Windows access opened on September 14, three weeks after the same thing arrived on Linux PCs and NVIDIA's DGX Spark desktop.

The model doing the work locally is Qwen 3.8 27B, which is where the memory requirement comes from. Pro and Max subscribers download the app from the Microsoft Store, and anything the agent finishes on the device uses none of the Computer credits a subscription buys.

Perplexity describes the agent analysing data, drawing information together across many files and handling recurring jobs, with harder reasoning passed up to cloud models when the person chooses to pass it. That covers most of the work firms have kept away from AI because the material is confidential. The contract folder, the month's reconciliations, the personnel file.

NVIDIA announced the Windows release on its own blog. Nobody gets the local privacy without buying the card first, and the cheapest route in, by The New Stack's reckoning this week, is an older RTX 3090 with 24 gigabytes, which still costs well over $1,500. A DGX Spark is $4,800.


In June, Anthropic started keeping 30 days of prompts and outputs from its most capable models, Fable 5 and Mythos 5, saying it needed the logs to catch attacks that play out across many requests and accounts. It said the data would not be used for training.

Three large customers found that insufficient. Nvidia restricts Claude to less sensitive internal work and prefers its own Nemotron models for proprietary jobs. Palantir has pressed for a guarantee of zero retention before it will offer the models inside its own software. Reuters reported those positions after The Information first described them.

A promise to store nothing is narrower than it sounds: the lab can still learn from how a customer uses the tool.

Anthropic has since replaced the policy with Enterprise Frontier Safeguards, which lets a customer keep those 30 days of logs in its own Amazon S3, Azure or Google Cloud storage, under its own keys, scanned automatically with no review by Anthropic staff. Palantir sells Claude on to its own customers and Nvidia makes the chips the labs run on; a company without that kind of leverage works under the terms printed in the plan it pays for.


Ask an automated phone line to look something up and it goes quiet while it does. Google's two new voice models talk through that pause: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking run their tool calls in the background while the conversation continues, saying out loud that the request has landed and narrating progress until the task finishes. They also hear when a caller switches language and switch with them, across 97 languages, in the middle of a call.

Both went out on September 15 through the Gemini API, with business access through Gemini Enterprise. Artificial Analysis priced 3.8 Live at $0.84 for an hour of input audio.

Sierra's banking leaderboard is the harder measure of the two Google cites, because it counts whether the agent resolves the customer's request. Extended Thinking leads it with 35.1%, ahead of GPT-Live-1 Astra at 32.0% and xAI-Realtime at 16.5%. The best voice agent anyone can buy this week finishes about one banking request in three.


Any company can buy the same general AI model as its competitor, at the same published price. The scarce input has moved one step back, to the people who can tell a model when it is wrong.

A grader is paid for the correction: the case where a confident answer is wrong, and the reason a professional rejects it. Every firm makes that material daily, in the recommendation an underwriter overrules and the exception a scheduler makes.

The startups that pay doctors, lawyers and scientists to produce corrections at scale have built large businesses on that work alone. Mercor and Handshake each passed a billion dollars in annual revenue this year.


Automated scanners are searching the internet for weak spots faster than governments and companies can patch systems that are decades old. Abigail Bradshaw, director-general of the Australian Signals Directorate, said so at the Sydney Dialogue on Monday, and nothing about the scanning depends on which country a server is in.

Old software is the part of this a manager can name without help: the finance package whose vendor stopped issuing fixes, the machine controller that cannot be updated without stopping the line. Both were manageable when an attacker had to find them by hand.

Bradshaw's proposal is an "early warning system" for AI risks, built by security agencies and AI companies together, modelled on the one that exists for cyber threats.


Salesforce's customer support team has gone from 9,000 people to about 5,000. Marc Benioff gave those figures himself, with the arithmetic behind them: roughly a million and a half customer conversations handled by agents over a period in which human agents handled a million and a half, satisfaction scores about the same, and support costs down 17% since the start of 2025 by Fortune's account. This is the company that spent Dreamforce this week selling the same agents to everyone else.

Salesforce now builds the model underneath those agents. Koa, Salesforce's first reasoning model, was post-trained with Nvidia on Nvidia's open-weight Nemotron using synthetic customer data. Kari Ann Briski of Nvidia told TechCrunch the pair built it for token efficiency, which means fewer tokens burned per task than sending the same job to Claude or ChatGPT. Claude still arrives through the new Claudeforce partnership.

The invoice has been harder to predict than the headcount. By Salesforce Ben's account, Agentforce launched in 2024 at $2 per conversation, added Flex Credits in May 2025, and gained pay-as-you-go and pre-commit options that August.


THE DAILY BRIEF

Get the next edition in your inbox.

A five-minute read, every weekday morning.

Free forever · Unsubscribe anytime

Share This Brief

Other Editions

Newest first


THE DAILY BRIEF

Read the next one first.

A five-minute read, every weekday morning.

Free forever · Unsubscribe anytime