Skip to main content
Aggregate PYMNTS 金融科技 15 Aug 2026 - 06:01

Speed Becomes the Product as OpenAI and Google Sell Faster AI

RSS 官方收录 · 可信分层展示

关键摘要

OpenAI and Google both released faster artificial intelligence models this week.…

  • Both put speed at the center of the pitch, and both are starting to tr…
  • OpenAI’s new tier is called Ultrafast, and for now it is a preview ope…
  • It runs the company’s GPT-5.

摘要引擎:抽取

正文提要

OpenAI and Google both released faster artificial intelligence models this week. Both put speed at the center of the pitch, and both are starting to treat response time as something businesses will pay for on its own.

OpenAI’s new tier is called Ultrafast, and for now it is a preview open to a small group of customers. It runs the company’s GPT-5.6 Sol model up to 14 times faster than the standard tier, at up to 750 output tokens per second, according to a Thursday (Aug. 13) company announcement. The model itself underneath is the same. It just answers faster, running on chips from a company called Cerebras.

Early customers include Jane Street, Podium, Basis and Rogo, which are testing it for coding, financial research, customer support, voice and commerce, per the announcement.

“Until now, getting real-time speed typically meant choosing a smaller or more specialized model,” the announcement said. “Ultrafast points to progress in a new direction: more useful work per second.”

Google launched its own faster model, Gemini 3.7 Flash, the same day. It costs $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31, half what the previous Flash model cost, then doubles on Jan. 1, 2027, to the price Gemini 3.6 Flash carried all along, according to a Thursday company blog post.

Google is also claiming a capability gain, calling 3.7 Flash its “most intelligent workhorse model yet for coding and agents” in the post.

Benchmarking firm Artificial Analysis clocked Gemini 3.7 Flash’s output at about 340 tokens per second, nearly three times the speed of GPT-5.6 Terra and GLM-5.2, and placed it on the frontier of intelligence versus time per task.

Both launches point to the same shift. For two years, AI pricing has mostly come down to how capable a model is and how much a company uses it. Speed is becoming a third feature companies are being asked to pay for on its own.

Speed Matters More for Fraud Checks Than for Overnight Reports

Not every AI task needs to be fast. A bank checking whether a transaction is fraudulent must decide in a fraction of a second. AI-driven fraud detection has already saved at least $5 million for 42% of card issuers, according to the PYMNTS Intelligence report “Where Payment Decisions Happen: How Issuer Data Is Powering the Next Era of Commerce.” A slow fraud check is not just annoying. It can mean letting a fraudulent charge go through before the system catches up.

In its announcement, OpenAI named voice applications, customer support, commerce, coding and financial research as the kind of work its fast tier is built for, along with incident response, which covers fraud and security threats that need a fast decision.

Compare that to a company running an AI job overnight to sort through old documents. Nobody is waiting on the other end. That job can run on cheaper, slower computers with no real cost to the business. It’s the same AI capability, but two different prices, depending only on whether someone is waiting on the answer.

AI Pricing Is Splitting Into a Fast Lane and a Slow One

This is similar to how companies already buy internet service by paying more for a guaranteed fast connection when it matters, and using a cheaper, slower connection everywhere else. Cerebras, the company powering OpenAI’s fast tier, made the same argument in its Thursday announcement, comparing the OpenAI launch to earlier tech shifts like the move from dial-up internet to broadband.

If that comparison holds, businesses will start splitting their AI spending into two lanes. Fast, expensive AI will be used for jobs where a delay costs real money, like a customer chatting live with a company or a fraud check happening in real time. Slow, cheaper AI will be used for jobs where nobody notices the wait, like an overnight report or a batch of paperwork.

This week’s launches from OpenAI and Google suggest both companies expect that split to become normal, something businesses will soon plan for deliberately rather than treat as an afterthought.

For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.

The post Speed Becomes the Product as OpenAI and Google Sell Faster AI appeared first on PYMNTS.com.

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表