Product Launch2 min read

Celeris Launches AI Model 15 Times Faster Than GPT-5-mini

By , Senior AI ConsultantPublished

Celeris Labs launched Celeris-1, an AI model that answers about 15 times faster than similar sized rivals by building its whole response at once instead of word by word, trading a small amount of accuracy for speed that could cut costs for voice bots and customer service tools.

Celeris Labs, a small AI research group founded by Tom Hamer and Jesse Clark, launched its first model this week, called Celeris-1. The pitch is simple: near-frontier intelligence at a fraction of the wait time. It answers in about 157 milliseconds, which the company says is roughly 15 times faster than a similarly sized model from OpenAI called GPT-5-mini, and even faster against the full GPT-5.

The trick is how the model actually writes its answer. Almost every chatbot on the market today, including ChatGPT, Claude, and Gemini, writes one word at a time, always glancing back at everything it already wrote before choosing the next word. That is why you watch text appear on screen word by word. Celeris-1 works differently. It sketches a rough version of the entire answer in one shot, then cleans it up in a handful of quick passes, more like a sculptor roughing out a whole figure before smoothing every part together. This approach is called diffusion, and because the model is not stuck writing one word before starting the next, it can produce far more text per second.

The accuracy cost is real but modest. On a common test of reasoning and knowledge, Celeris-1 scored a few points below GPT-5 and GPT-5-mini. For tasks where the right answer truly matters, like drafting a legal clause or diagnosing a technical problem, that gap is worth caring about. For high-volume repetitive work, like sorting support tickets, pulling data out of forms, or running a voice assistant that needs to respond the instant a customer stops talking, the tradeoff usually favors speed.

There is a real limitation worth knowing before adopting it. Celeris-1 can only take in a few thousand words at a time, far less than most modern models. That rules it out for summarizing long contracts, lengthy call transcripts, or big spreadsheets in one pass. It is priced at 2 dollars per million words of input and 6 dollars per million words of output, and any business can sign up for the API today and plug it into existing software without new hardware.

What makes this launch worth watching is not really Celeris-1 itself, but what it signals. This is only the second commercially sold model built this way, following a similar product from a company called Inception Labs. The founders previously ran a search startup before pivoting to build a foundation model from scratch, a jump that costs enormous money and rarely works. Their willingness to make that bet suggests real investor appetite is forming around speed as a selling point on its own, separate from raw intelligence.

For most companies, this is not an urgent decision. But any business running real-time voice support, live chat, or high-volume document sorting should watch this space. The lesson is not to switch models based on a press release. It is to recognize that speed is becoming its own competitive category in AI, and the fastest option is not always the biggest name you have heard of.

Stay informed

Get AI intelligence like this delivered to your inbox.


You May Also Find Valuable