Every time your staff asks an AI tool a question or lets an automated process summarize a document, the company on the other side pays a computing bill. That bill is now the central financial problem in the AI industry. Anthropic spent an estimated $6.8 billion on computing in 2025 alone, across training and running its models. The cost of running AI at scale is not a background detail; it is the thing that determines whether AI businesses are sustainable.
Etched, a San Jose startup founded in 2022 by three Harvard dropouts, has built a chip called Sohu designed specifically to reduce that cost. The chip was manufactured by TSMC, the world's leading chipmaker, and is now in testing with early customers. Etched has already booked $1 billion in orders for full systems built around it.
The funding numbers tell their own story. The company raised $800 million in total, including a $500 million round closed last December that valued it at $5 billion. The investor list includes trading firms Jane Street and Two Sigma, financier Stanley Druckenmiller, Peter Thiel, and a group of prominent AI researchers. Three years ago, when the founders pitched the same idea to investors, every major firm passed. The company was reportedly running out of cash by the month.
What changed is that the AI industry moved from building models to running them at scale, and running them at scale is now brutally expensive. Gartner estimates that agentic AI workflows, where an AI agent handles a multi-step task autonomously, require between 5 and 30 times more computing per task than a simple chatbot query. Enterprises that sized their AI budgets on pilot data have received production bills that bear no resemblance to what they modeled.
Etched's answer is specialization. Sohu is hardwired to run only one type of AI model architecture, the transformer architecture that underlies almost every major AI model in commercial use today, including ChatGPT, Claude, and Gemini. Because the chip does nothing else, every component of the silicon is optimized for that one job. The company claims an eight-chip Sohu server generates more than 500,000 responses per second on a standard benchmark, compared to roughly 23,000 for an equivalent Nvidia server. None of those claims have been independently verified yet.
The risk is equally clear. If the AI industry migrates to a different model architecture, Sohu becomes unusable. The chip cannot run other types of AI, and that limitation is built into the hardware permanently. Etched's CEO has acknowledged this openly: the company is betting that transformers remain dominant. So far that bet has held, but AI research moves fast and no one can guarantee it.
For business operators, the relevant takeaway is not about this chip specifically. It is about the market signal: the AI inference market was valued at roughly $106 billion in 2025 and is growing at nearly 20% annually. That growth is being driven by companies like yours consuming AI at higher volumes, which means the cost of that consumption is rising too. Cheaper, more efficient chips from companies like Etched, Groq, and Cerebras could eventually lower the per-query cost you pay through AI service providers. But that benefit only reaches you if the providers pass it along.
The more immediate question for any organization running AI at volume is whether its current cost model still holds. Pilots are cheap. Production, at real scale, is not. The businesses building discipline around how they measure and control AI running costs now will be in a much stronger position than those treating it as a line item to revisit later.