Microsoft has begun replacing select OpenAI and Anthropic frontier models with its own MAI model family in production workloads inside Excel and Outlook. Tens of thousands of weekly prompts previously handled by third-party models are now processed by Microsoft's internal AI systems.
The move reflects growing pressure on hyperscalers to control inference costs amid massive AI spending. While still a small fraction of overall Copilot traffic, it demonstrates Microsoft's accelerating independence from external providers in core productivity tools.
Last updated August 4, 2026: added real benchmark results published since the original report, including MAI-Cyber-1-Flash beating Anthropic and OpenAI on a security benchmark, and Microsoft's more public competitive posture (sales training, earnings-call framing) against its own AI partners.
What Happened: Microsoft Shifts to In-House MAI Models
Bloomberg reported on July 7, 2026 that Microsoft has shifted some Excel and Outlook AI workloads from OpenAI and Anthropic models to its own MAI family. The change affects tens of thousands of weekly prompts, representing a meaningful volume of production traffic. The routing is task-based: routine, high-volume requests such as summaries, drafting and formula help go to MAI, while frontier-grade tasks like complex reasoning, novel research and agentic workflows still route to OpenAI or Anthropic models.
This follows Microsoft's Build 2026 developer conference in June, where the company unveiled seven new MAI models, including MAI-Thinking-1 for complex reasoning alongside models for coding, image generation, speech and transcription, all trained from scratch on licensed data. Microsoft AI chief Mustafa Suleyman said in June that the company is trying to reduce, and eventually eliminate, what it pays Anthropic for inference: "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost."
Key Details: Real Benchmarks, Not Just Cost Claims
Since the July 7 Excel/Outlook shift, Microsoft has published more concrete evidence its in-house models can actually compete. MAI-Thinking-1, the flagship reasoning model unveiled at Build 2026, runs 35 billion parameters with a 256K context window and, Microsoft says, matches or outperforms GPT-5.5 and Claude Opus 4.6 on SWE-Bench Pro coding benchmarks, claiming up to 10x better cost efficiency in tuned enterprise workloads. It's not winning everywhere: MAI models reportedly still fall short of Google's Gemini 2.0 Ultra and Anthropic's Claude 3.5 Sonnet on complex reasoning and multilingual tasks. More strikingly, Microsoft announced on July 27 that its MAI-Cyber-1-Flash security model scored 95.95% on the CyberGym vulnerability benchmark, beating both Anthropic's and OpenAI's models by 12 points, at half the inference cost.
The shift extends beyond product engineering into sales: Microsoft has reportedly begun training its own sales staff to actively promote in-house AI over OpenAI and Anthropic to customers, and executives have started openly describing Microsoft as a direct rival to both companies on recent earnings calls, a more confrontational public posture than the "cost optimization" framing suggests on its own.
The task-based routing strategy Microsoft is using, sending routine requests to MAI while keeping frontier tasks on OpenAI and Anthropic, reflects a broader industry pattern often called a "model router" or "mixture of providers" approach: rather than committing all traffic to a single model regardless of task complexity, a system evaluates what a given request actually requires and sends it to the cheapest model capable of handling it well. Summarizing a spreadsheet or drafting a routine email doesn't require the same reasoning depth as complex multi-step analysis, so routing the former to a smaller, cheaper in-house model while reserving expensive frontier-model calls for tasks that genuinely need them is a straightforward cost optimization once a company has a competitive in-house alternative for at least the simpler tier of requests. What makes Microsoft's version notable is the scale at which it can deploy this: few companies have Excel and Outlook's install base to generate the kind of routine-task volume that makes building and running an in-house model economically worthwhile in the first place.
Why It Matters
This is no longer just about routing routine Excel formulas to a cheaper internal model. A cybersecurity model beating Anthropic and OpenAI on an independent benchmark, sales teams pitching against Microsoft's own AI partners, and executives calling Microsoft a rival on earnings calls together describe a company hedging its multi-billion-dollar OpenAI partnership in a much more public way than "we're saving on inference costs" implies. For enterprises, it signals that exclusive AI vendor relationships, even Microsoft's famously deep one with OpenAI, aren't permanent once a company has the resources to build credible alternatives.
What Happens Next
Microsoft will likely keep expanding MAI usage into workloads where it's shown competitive benchmark results (cybersecurity, coding) while continuing to route complex reasoning and multilingual tasks to Anthropic and OpenAI models where MAI still lags. Whether Microsoft's sales-side push against its own AI partners changes the terms of its OpenAI relationship, rather than just its internal cost structure, is the bigger strategic story to watch.
Final Takeaway
What started as quiet cost-optimization in Excel and Outlook has become a more visible competitive posture: real benchmark wins in specific domains (cybersecurity, coding), acknowledged gaps in others (complex reasoning, multilingual), and a sales and messaging strategy that now openly positions Microsoft against the AI labs it also depends on.
Key Points
- Microsoft's MAI-Cyber-1-Flash scored 95.95% on the CyberGym benchmark, beating Anthropic and OpenAI models by 12 points at half the inference cost, announced July 27, 2026.
- MAI-Thinking-1 reportedly matches or beats GPT-5.5 and Claude Opus 4.6 on coding benchmarks, but still trails Gemini 2.0 Ultra and Claude 3.5 Sonnet on complex reasoning and multilingual tasks.
- Microsoft has reportedly begun training sales staff to promote in-house AI over OpenAI and Anthropic, and executives now describe Microsoft as a direct rival to both on earnings calls.
Microsoft's MAI Model Family and Task Routing
The seven MAI models Microsoft introduced at Build 2026 span reasoning, coding, image generation, speech and transcription, with MAI-Thinking-1 positioned as the flagship reasoning model. Microsoft has said one of its coding-focused MAI models can match the capabilities of Anthropic's Opus 4.6 on certain benchmarks at a lower cost, part of the rationale for shifting workloads in-house.
The routing approach is deliberately selective rather than a wholesale replacement: high-volume, routine requests inside Excel and Outlook, such as summarizing a spreadsheet or drafting a reply, are well suited to MAI's current capabilities, while more demanding, frontier-grade tasks continue to route to OpenAI or Anthropic models. Mustafa Suleyman's stated goal is to reduce, and eventually eliminate, what Microsoft pays Anthropic for inference. Microsoft's underlying multi-billion-dollar OpenAI contract remains in place, though the company's public posture toward OpenAI has grown more openly competitive since this shift began.
FAQs
Sources and Verification
- Bloomberg, July 7, 2026
- FourWeekMBA analysis, July 2026
- Tech Times: MAI-Cyber-1-Flash benchmark results, August 2026
- TechCrunch: Microsoft's competitive posture and sales strategy
This article was reviewed as part of CapisTech's editorial fact-checking process.



