Skip to main content
Cloud / Azure / Products / Phi Open Models - Microsoft Small Language Models

Phi Open Models - Microsoft Small Language Models

Microsoft Phi Open Models are compact, efficient language models, currently including Phi-4 and Phi-4-reasoning variants.

ai-machine-learning
Pricing Model Open weights free to download; hosted via Microsoft Foundry billed per token or per compute hour
Availability Available in Microsoft Foundry
Data Sovereignty EU regions available
Reliability SLA as published by the provider SLA

What are Microsoft Phi Models?

Microsoft Phi is a family of small language models (SLMs) that deliver surprisingly high performance at compact sizes. Unlike large LLMs, Phi models are optimized for scenarios where resource efficiency, latency, or privacy matter more than maximum all-round capability.

The current generation includes Phi-4 and derived reasoning variants such as Phi-4-reasoning, Phi-4-reasoning-plus, Phi-4-mini-reasoning, and the latency-optimized Phi-4-mini-flash-reasoning for edge scenarios, plus Phi-4-reasoning-vision for multimodal tasks. The models are available as open weights via Hugging Face, in the Microsoft Foundry model catalog (formerly Azure AI Studio / Azure AI Foundry), and as locally runnable variants via Microsoft Foundry Local.

The models are well suited for edge deployment, mobile apps, or scenarios with limited GPU capacity that still require a compact but capable model.

Core Features

  • Compact model sizes ranging from a few billion parameters (e.g. Phi-4-mini) up to Phi-4 with roughly 14 billion parameters
  • Specialized reasoning variants for math, multi-step problem solving, and agentic use cases
  • Available in the Microsoft Foundry model catalog and as open weights
  • Local execution via ONNX exports and Microsoft Foundry Local
  • Multimodal variants for text, image, and visual reasoning

Typical Use Cases

Edge AI: Deploying language models on resource-constrained devices such as IoT gateways or local servers.

Mobile Applications: AI features in apps without a cloud roundtrip for better latency and offline capability.

Reasoning-Intensive Tasks: Mathematical problem solving, tutoring applications, and agentic workflows using the Phi-4-reasoning variants at a lower resource footprint than large foundation models.

Benefits

  • Significantly lower inference costs than large LLMs
  • Faster response times through compact size
  • Local execution possible for privacy and offline operation
  • Open weights available for customization and self-hosting

Frequently Asked Questions

How do Phi models differ from large foundation models?

Phi models are considerably smaller and more resource-efficient, but tend to be less versatile for very broad or highly complex tasks. They are particularly well suited for focused use cases such as reasoning, edge deployment, or mobile scenarios.

Can I run Phi models locally?

Yes, Phi models are available as open weights and ONNX exports and can run via Microsoft Foundry Local or your own infrastructure without a cloud connection.

What does using Phi models cost?

The model weights can be downloaded and self-hosted free of charge. If you use the models as a managed service via Microsoft Foundry, billing depends on the deployment type: per token usage or per compute hour for dedicated endpoints.

Which model variants are currently available?

The Phi family currently includes Phi-4, the reasoning variants Phi-4-reasoning, Phi-4-reasoning-plus, and Phi-4-mini-reasoning, the latency-optimized Phi-4-mini-flash-reasoning, and Phi-4-reasoning-vision for multimodal tasks. The exact catalog selection changes as new releases ship.

How do I use Phi in Azure?

Phi models are available via the Microsoft Foundry model catalog and can be deployed as a serverless endpoint billed per token or as a managed compute deployment.

Integration with innFactory

As a Microsoft Solutions Partner, innFactory supports you with Phi models: evaluation for your use case, fine-tuning, edge deployment, and integration into existing applications.

Typical Use Cases

Edge AI applications
Resource-efficient inference
Mobile AI solutions
On-device language models

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Azure (official documentation). This page does not represent an offer by Azure.

Microsoft Solutions Partner

innFactory is a Microsoft Solutions Partner. We provide expert consulting, implementation, and managed services for Azure.

Microsoft Solutions Partner Microsoft Data & AI

Similar Products from Other Clouds

Other cloud providers offer comparable services in this category. As a multi-cloud partner, we help you choose the right solution.

Google Cloud

Agent Platform Workbench (formerly Vertex AI Workbench)

Agent Platform Workbench provides managed JupyterLab environments on the Gemini Enterprise Agent Platform with BigQuery, …

Pricing Usage-based via the Gemini Enterprise …
SLA SLA as published by the provider
Compare →
Google Cloud

Cloud Talent Solution - Job Search with Machine Learning

Cloud Talent Solution brings machine learning to the job search experience and supports recruiting platforms with job …

Pricing Pricing as of January 5, 2021: charged …
SLA As published by the provider / see official documentation
Compare →
Google Cloud

Deep Learning VM Images - Preconfigured ML VM Images

Deep Learning VM Images are virtual machine images optimized for data science and machine learning with frameworks such …

Pricing The images themselves are free to use …
SLA SLA as published by the provider
Compare →
Google Cloud

Enterprise Knowledge Graph - Entity Reconciliation for BigQuery

Enterprise Knowledge Graph reconciles siloed data through an entity reconciliation API and adds lookups via the Google …

Pricing Google does not publish a dedicated …
SLA Pre-GA products are available "as is" and might have limited support per the documentation; As published by the provider / see official documentation
Compare →
Google Cloud

Genkit - Open Source Framework for AI Applications

Genkit is Google's open-source framework for building full-stack, AI-powered and agentic applications with TypeScript, …

Pricing Open-source framework; costs arise from …
SLA No SLA applies to the framework itself; SLAs apply to the services used as published by their providers
Compare →
Google Cloud

Model Garden - Model Catalog of the Agent Platform

Model Garden is the model library of the Gemini Enterprise Agent Platform for discovering, testing, customizing, and …

Pricing For open source models you are charged …
SLA SLA as published by the provider
Compare →

89 comparable products found across other clouds.

Ready to start with Phi Open Models - Microsoft Small Language Models?

Our certified Azure experts help you with architecture, integration, and optimization.

Schedule Consultation