Microsoft’s MAI models cut GPU costs up to 89% as Bing and Dynamics go in-house
New public preview releases plus production metrics are Microsoft’s most aggressive case yet to shrink reliance on OpenAI.

Microsoft’s Superintelligence team released MAI-Image-2.5-Pro and MAI-Voice-2-Flash into public preview, alongside production data showing major GPU cost reductions. For execs, it signals a shift from “AI demos” to controllable, first-party infrastructure across Bing, PowerPoint, OneDrive, Dynamics 365, and Azure.
On Wednesday, Microsoft did two things at once: it shipped new in-house AI models into public preview, and it published production numbers meant to prove those models can run real products without leaning as hard on OpenAI. The GPU-cost headline is the one that makes CFOs sit up: Microsoft says its MAI-Image-2.5 and MAI-Voice-2-Flash cut GPU costs by up to 89% versus OpenAI models in specific deployments.
The stakes are clearer when you look at where Microsoft says the models now live. Bing Image Creator runs entirely on MAI-Image-2.5 end to end. PowerPoint, OneDrive, Dynamics 365 Contact Center, and Azure Voice Live also move pieces of core workflows onto Microsoft’s own models. Microsoft CEO Satya Nadella frames it as “route traffic” to MAI when Microsoft’s models match or outperform frontier alternatives, while still keeping OpenAI and Anthropic models in the “orchestration system.” The subtext is simple: Microsoft wants to be the one deciding the bill, not the one paying it.
Let’s start with the two new public preview releases, because they define the strategy at both ends of the quality-speed tradeoff. MAI-Image-2.5-Pro is pitched as Microsoft’s highest-fidelity image generator to date, with pricing tied to input and output token counts. It costs $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. Microsoft also says the model targets the premium tier, including detailed editing and precise in-image text rendering, which it calls out as a long-standing weakness for image generation models. On the community side, it notes MAI-Image-2.5 recently launched at No. 2 for image editing on Arena, a leaderboard community view that has become a de facto generative media scoreboard.
MAI-Voice-2-Flash is the opposite style of bet. It was first previewed at Microsoft’s Build conference and is designed for high-volume enterprise voice workloads where latency and cost per call dominate. Microsoft says Flash runs twice as fast as MAI-Voice-2 and costs 32% less, priced at $15 per million characters. Put together, these launches are not one monolithic flagship. They are model families built for different product realities, from creative studios chasing fidelity to contact centers that care about throughput.
Now the “prove it in production” part. Microsoft’s announcement includes an unusual level of specificity about deployments and measured outcomes. For image generation, it claims Bing Image Creator runs entirely on MAI-Image-2.5. In PowerPoint, Microsoft says MAI-Image-2.5 reduces GPU costs by up to 84% compared with GPT-Image-2, OpenAI’s image model. In OneDrive, where MAI-Image-2.5 is the default for key image-editing scenarios, Microsoft reports a 26% increase in save rates, roughly 25% lower P95 latency, and 2.5 times greater efficiency under medium-utilization production workloads.
On the voice side, Microsoft says MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, used by customers including T-Mobile and EasyJet, and claims GPU cost reductions of up to 89%. It also integrates into Azure Voice Live for developers building speech-to-speech agents. The most consequential deployment, at least by consequence to humans, is healthcare. Microsoft’s Dragon Copilot, used by 170,000 medical providers and responsible for processing 28 million patient encounters last quarter, now runs on MAI-Transcribe-1.5 for its multilingual workflow across 58 languages. Microsoft claims internal evaluations show a 50% relative reduction in both transcription and language-identification error rates across most languages. In clinical notes, smaller errors can cascade, so “relative reduction” matters even if it doesn’t sound flashy.
If you’re wondering how Microsoft’s models get good without instantly needing the newest GPUs, the companion post it published the same day helps explain the mechanism. Microsoft calls its approach a “hill-climbing machine,” an integrated flywheel of data, models, and the product “harness.” The clearest example is MAI-Code-1-Flash, launched in GitHub Copilot in June. Microsoft claims it achieves approximately 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, while using 10% fewer median tokens. It also says users are 6% more likely to return across multiple days versus GPT-5.4 Mini, and 11% more likely than with Claude Haiku 4.5.
Then Microsoft goes further: it takes the MAI-Code-1-Flash checkpoint and further trains it inside an Excel reinforcement learning environment. The goal is to teach a coding model the tools and workflows of spreadsheet knowledge work. Microsoft says, based on production user feedback, it lands on par with GPT-5.6 for the most common Excel tasks, while being small enough to run on Nvidia’s older H100 and even A100 GPUs instead of requiring the latest-generation accelerators. Executives should treat that hardware detail as more than engineering trivia. In a market where chip allocation can become a bottleneck, the ability to serve “frontier-adjacent quality” on two-generation-old silicon shifts the economics of deployment. It also frees newer hardware, including Microsoft’s now-operational GB200 cluster, for training rather than serving.
Finally, the relationship strategy. Nadella’s post on X, “Frontier Diffusion & Control,” reads like a roadmap for avoiding single-vendor dependency. Microsoft says it can “take saturated frontier capabilities and deliver them at scale and at lower cost through models optimized for high-usage products,” while continuing to use frontier models for “frontier needs.” It also says it is “beginning to route traffic across our first-party surfaces to MAI whenever our models match or outperform frontier alternatives.” At the same time, Nadella stresses that frontier models from OpenAI and Anthropic are part of the orchestration system alongside MAI.
This matters because Reuters reported in April that Microsoft’s exclusive license to OpenAI’s technology was revised into a non-exclusive arrangement, and The Information reported last September that Microsoft had begun incorporating Anthropic models into some products. Wednesday’s announcement completes the picture Microsoft wants: Microsoft as orchestrator, with partners’ frontier models as interchangeable components, and Microsoft models absorbing a growing share of routine traffic. For boards and leadership teams in AI-heavy businesses, the question stops being “who has the best model?” and becomes “who controls cost, routing, and infrastructure at scale when millions of users show up every day?”
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Gatekeeper lets user-run macOS apps be silently swapped after first launch
Researchers show how “signed” internet downloads can become doppelgangers without reauthorization, forcing a rethink of macOS trust.

Alphabet’s $800B commitments and cash-flow flip spook Wall Street, dragging Mag 7 lower
Alphabet hits first-ever negative cash flow, warns 2027 capex is higher, and investors punish AI spending they can’t ignore.

Oracle locks a 10-year Pentagon software deal worth up to $7B
The on-premises commitment reshapes DoD IT procurement, and the ripple reaches every enterprise vendor pitching Washington.

