Stories

Everyday AI search tools push whole teams toward the same handful of ideas

Researchers took over five AI browser assistants with one email, and the best AI agent finished 26.2% of 1,490 real work assignments in a Berkeley test.

By , Senior AI ConsultantEdition of

6stories
4minute read
In This Edition

Two hundred and forty-five people around the world entered an online challenge to find new ways of cutting food waste. Some researched with an ordinary search tool. The rest used a version the researchers had rebuilt to return unusual, less popular material, and among their ideas the researchers counted twice as many distinct groups.

Ordinary search engines and chatbots put the most popular, highest-probability answer at the top. That ranking makes them fast when someone needs a known answer, and it means twelve colleagues asking the same question get nearly the same twelve answers. The study, published in the Academy of Management Journal by Moran Lazar of Tel Aviv University, Hila Lifshitz of Warwick Business School and two co-authors, calls those clusters of near-identical thinking ideation bubbles.

The researchers list two changes to how the question is put: asking for approaches used in other industries, and asking for several framings of a problem instead of one best answer.

An unfamiliar comparison is only useful to someone who can see what it connects to, and in this study the experts got most of the benefit from the rebuilt tool, which runs against the common claim that AI raises a beginner to an expert's level.


At Black Hat in Las Vegas earlier this month, researchers from the security firm Zenity Labs took over five AI browser assistants using nothing but content the assistant read: an email, a calendar invitation, a link under a social media post. The person under attack clicked nothing. The five were Claude in Chrome, Gemini in Chrome, Perplexity's Comet, OpenAI's Atlas and Copilot in Edge.

Zenity calls the flaw intent collision. The assistant takes in the instruction from its user and the text of whatever it reads through the same channel, and it cannot reliably tell them apart, so a sentence buried in an invitation is carried out the same way as an instruction typed by its owner.

In one chain, Atlas was sent to Amazon, where it filled a cart and set the delivery address to the attacker's; blocked from pressing the purchase button itself, it asked Amazon's own shopping assistant, Rufus, to place the order.

Zenity reported its findings to OpenAI and Anthropic in late 2025 and early 2026; some were fixed, and some vendors answered that the behavior is intended, because reading a page and acting on it across the sites where a user is signed in is what these assistants were built to do.

OpenAI shut Atlas down on 9 August and moved its browser-agent features into ChatGPT and Codex.


Salesforce's own research says its customers now run three times as many AI agents on Agentforce as a year ago. Upwork's chief executive, Hayden Brown, says 23% of businesses that moved work to AI have moved it back to people.

Outside analysts say the Salesforce number measures activity, how many agents are running and how much they attempt, and says nothing about whether the business is better off.

Upwork measured the other end. Brown says that even on the smallest jobs, the kind a company pays a couple of hundred dollars for, agents "complete those tasks seventy percent better if they have human oversight."

Researchers at the University of California, Berkeley collected 1,490 real work assignments from more than 250 professionals across 55 industries: preparing a legal filing, building a financial model, designing a manufacturing part, work that takes a person hours or weeks. The best system, OpenAI's Codex running on GPT-5.5, delivered 26.2% of them correctly. A missing deliverable or a wrong one counted as a failure, with no credit for a good attempt.


Two customers see different prices for the same seat. What produced the difference, the market or the buyer?

Virgin Atlantic prices seats with a system that repeats that calculation in real time. MIT Technology Review traced the tool back to Fetcherr, whose market model is spreading through the airline industry while US lawmakers investigate whether systems of this kind amount to surveillance pricing.

The two kinds of pricing look identical from the outside. A market model reads what any competitor could also see: seats unsold, days to departure, what other airlines are charging this morning. A surveillance pricing system reads the buyer: the device, the purchase record, an estimate of what this particular person will pay. The second is what Congress is asking about.

Delta's president, Glen Hauenstein, told investors in July 2025 that about 3% of the airline's domestic fares were being set by Fetcherr's system, and that Delta wanted that near 20% by the end of the year.


The person who has to explain an AI rollout to a team of twelve is their line manager.

Gallup asked HR chiefs how that is going. 99% of them call AI central to their company's strategy, and half say they do not trust their managers to guide employees through the change.

A company announcement can say the technology is central. It cannot tell a payroll clerk whether her Tuesday changes, and that answer has to come from the manager who runs her week. Gartner, ManpowerGroup and Mercer each measured the same gap this year, in separate studies of employers.


Adobe switched on three audio tools inside Firefly: generated music, generated voiceover and sound effects, free to use in the app and cleared for commercial work, with Adobe saying it will cover the legal costs if a customer's use is challenged.

Record labels are suing Suno and Udio, the two best known standalone AI music services. A team making a training video or an internal announcement can now generate the audio without paying for a stock music license or a voice actor.


THE DAILY BRIEF

Get the next edition in your inbox.

A five-minute read, every weekday morning.

Free forever · Unsubscribe anytime

Share This Brief