Skip to main content
Cloud / Azure / Products / Phi Open Models - Microsoft Small Language Models

Phi Open Models - Microsoft Small Language Models

Microsoft Phi Open Models are compact, efficient language models, currently including Phi-4 and Phi-4-reasoning variants.

ai-machine-learning
Pricing Model Open weights free to download; hosted via Microsoft Foundry billed per token or per compute hour
Availability Available in Microsoft Foundry
Data Sovereignty EU regions available
Reliability SLA as published by the provider SLA

What are Microsoft Phi Models?

Microsoft Phi is a family of small language models (SLMs) that deliver surprisingly high performance at compact sizes. Unlike large LLMs, Phi models are optimized for scenarios where resource efficiency, latency, or privacy matter more than maximum all-round capability.

The current generation includes Phi-4 and derived reasoning variants such as Phi-4-reasoning, Phi-4-reasoning-plus, Phi-4-mini-reasoning, and the latency-optimized Phi-4-mini-flash-reasoning for edge scenarios, plus Phi-4-reasoning-vision for multimodal tasks. The models are available as open weights via Hugging Face, in the Microsoft Foundry model catalog (formerly Azure AI Studio / Azure AI Foundry), and as locally runnable variants via Microsoft Foundry Local.

The models are well suited for edge deployment, mobile apps, or scenarios with limited GPU capacity that still require a compact but capable model.

Core Features

  • Compact model sizes ranging from a few billion parameters (e.g. Phi-4-mini) up to Phi-4 with roughly 14 billion parameters
  • Specialized reasoning variants for math, multi-step problem solving, and agentic use cases
  • Available in the Microsoft Foundry model catalog and as open weights
  • Local execution via ONNX exports and Microsoft Foundry Local
  • Multimodal variants for text, image, and visual reasoning

Typical Use Cases

Edge AI: Deploying language models on resource-constrained devices such as IoT gateways or local servers.

Mobile Applications: AI features in apps without a cloud roundtrip for better latency and offline capability.

Reasoning-Intensive Tasks: Mathematical problem solving, tutoring applications, and agentic workflows using the Phi-4-reasoning variants at a lower resource footprint than large foundation models.

Benefits

  • Significantly lower inference costs than large LLMs
  • Faster response times through compact size
  • Local execution possible for privacy and offline operation
  • Open weights available for customization and self-hosting

Frequently Asked Questions

How do Phi models differ from large foundation models?

Phi models are considerably smaller and more resource-efficient, but tend to be less versatile for very broad or highly complex tasks. They are particularly well suited for focused use cases such as reasoning, edge deployment, or mobile scenarios.

Can I run Phi models locally?

Yes, Phi models are available as open weights and ONNX exports and can run via Microsoft Foundry Local or your own infrastructure without a cloud connection.

What does using Phi models cost?

The model weights can be downloaded and self-hosted free of charge. If you use the models as a managed service via Microsoft Foundry, billing depends on the deployment type: per token usage or per compute hour for dedicated endpoints.

Which model variants are currently available?

The Phi family currently includes Phi-4, the reasoning variants Phi-4-reasoning, Phi-4-reasoning-plus, and Phi-4-mini-reasoning, the latency-optimized Phi-4-mini-flash-reasoning, and Phi-4-reasoning-vision for multimodal tasks. The exact catalog selection changes as new releases ship.

How do I use Phi in Azure?

Phi models are available via the Microsoft Foundry model catalog and can be deployed as a serverless endpoint billed per token or as a managed compute deployment.

Integration with innFactory

As a Microsoft Solutions Partner, innFactory supports you with Phi models: evaluation for your use case, fine-tuning, edge deployment, and integration into existing applications.

Typical Use Cases

Edge AI applications
Resource-efficient inference
Mobile AI solutions
On-device language models

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Azure (official documentation). This page does not represent an offer by Azure.

Microsoft Solutions Partner

innFactory is a Microsoft Solutions Partner. We provide expert consulting, implementation, and managed services for Azure.

Microsoft Solutions Partner Microsoft Data & AI

Similar Products from Other Clouds

Other cloud providers offer comparable services in this category. As a multi-cloud partner, we help you choose the right solution.

STACKIT

STACKIT AI Model Experiments: Managed MLflow

STACKIT AI Model Experiments: managed MLflow for experiment tracking, LLM tracing, and EU AI Act audit trails on EU …

Pricing Public Preview: service itself free, …
Compare →
STACKIT

STACKIT AI Model Serving: Sovereign LLMs

STACKIT AI Model Serving: Run open-weight LLMs like Llama, Qwen, and GPT-OSS GDPR-compliant from German data centers, …

Pricing Pay-as-you-go per token (input/output)
SLA Runs on the data-sovereign STACKIT Cloud
Compare →
STACKIT

STACKIT Dremio - Sovereign Data Lakehouse

STACKIT Dremio: managed data lakehouse based on Dremio for SQL queries without data movement. GDPR-compliant, public …

Pricing Public preview; unified billing via …
SLA SLA as published by the provider; the service is currently in public preview
Compare →
STACKIT

STACKIT Intake - Data Ingestion for the Data Lakehouse

STACKIT Intake: managed data ingestion via the Kafka protocol directly into Apache Iceberg tables for the STACKIT Data …

Pricing Capacity-based; capacity is configured …
SLA SLA as published by the provider
Compare →
STACKIT

STACKIT Notebooks - Managed JupyterHub

STACKIT Notebooks: managed JupyterHub/JupyterLab for data science and ML in German and Austrian data centers.

Pricing Pay-per-use (compute resources of the …
SLA SLA as published by the provider
Compare →
STACKIT

STACKIT Workflows - Managed Apache Airflow

STACKIT Workflows is a managed workflow orchestration service built on Apache Airflow for data pipelines and ML …

Pricing Pay-per-use
SLA SLA as published by the provider
Compare →

89 comparable products found across other clouds.

Ready to start with Phi Open Models - Microsoft Small Language Models?

Our certified Azure experts help you with architecture, integration, and optimization.

Schedule Consultation