TLDR: Microsoft has launched Phi-4-mini-flash-reasoning, a new 3.8-billion-parameter AI model designed for high-speed, on-device logical reasoning. The model boasts up to 10 times the throughput and a two-to-threefold reduction in latency, making it ideal for mobile applications and edge devices. It introduces a novel ‘SambaY’ architecture and is available on Azure AI Foundry, Hugging Face, and the NVIDIA API Catalogue.
Microsoft has introduced a new addition to its Phi family of AI models, named Phi-4-mini-flash-reasoning. This compact, 3.8-billion-parameter model is engineered to deliver advanced reasoning capabilities in environments with limited computing power, memory, and speed, such as mobile apps and edge devices.
The company states that the new model can provide a tenfold increase in throughput and reduce average latency by two to three times compared to its predecessors. At the heart of this new model is a novel ‘decoder-hybrid-decoder’ architecture called SambaY. This architecture incorporates a Gated Memory Unit (GMU) that, combined with state-space models (Mamba) and sliding window attention, significantly reduces computational complexity and enhances performance on long-context tasks.
Phi-4-mini-flash-reasoning supports a 64k token context length and has been fine-tuned on high-quality synthetic data, with a particular focus on structured and mathematical reasoning tasks. According to benchmarks shared by Microsoft, the model outperforms AI models twice its size on tasks like AIME24/25 and Math500, while maintaining faster response times.
This makes it particularly well-suited for applications such as real-time tutoring tools, adaptive learning platforms, and mobile study aids.
In line with its commitment to responsible AI, Microsoft has integrated several safety measures into the model’s development. These include supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning from human feedback (RLHF). The company emphasizes that all Phi models are developed in accordance with its core principles of transparency, privacy, and inclusiveness.
Also Read:
- Microsoft’s AI-Driven Efficiency Yields Over $500 Million in Savings, Highlighting Ethical AI’s Business Value
- Microsoft Slashes Azure Generative AI Pricing by 60% for Multimedia Understanding
Developers can access Phi-4-mini-flash-reasoning through Azure AI Foundry, Hugging Face, and the NVIDIA API Catalogue. Its efficiency on a single GPU and its availability on major platforms are intended to allow for easy integration into existing developer workflows.


