Skip to main content
Cloud / Azure / Products / Phi Open Models - Microsoft Small Language Models

Phi Open Models - Microsoft Small Language Models

Microsoft Phi Open Models are compact, efficient language models, currently including Phi-4 and Phi-4-reasoning variants.

ai-machine-learning
Pricing Model Open weights free to download; hosted via Microsoft Foundry billed per token or per compute hour
Availability Available in Microsoft Foundry
Data Sovereignty EU regions available
Reliability SLA as published by the provider SLA

What are Microsoft Phi Models?

Microsoft Phi is a family of small language models (SLMs) that deliver surprisingly high performance at compact sizes. Unlike large LLMs, Phi models are optimized for scenarios where resource efficiency, latency, or privacy matter more than maximum all-round capability.

The current generation includes Phi-4 and derived reasoning variants such as Phi-4-reasoning, Phi-4-reasoning-plus, Phi-4-mini-reasoning, and the latency-optimized Phi-4-mini-flash-reasoning for edge scenarios, plus Phi-4-reasoning-vision for multimodal tasks. The models are available as open weights via Hugging Face, in the Microsoft Foundry model catalog (formerly Azure AI Studio / Azure AI Foundry), and as locally runnable variants via Microsoft Foundry Local.

The models are well suited for edge deployment, mobile apps, or scenarios with limited GPU capacity that still require a compact but capable model.

Core Features

  • Compact model sizes ranging from a few billion parameters (e.g. Phi-4-mini) up to Phi-4 with roughly 14 billion parameters
  • Specialized reasoning variants for math, multi-step problem solving, and agentic use cases
  • Available in the Microsoft Foundry model catalog and as open weights
  • Local execution via ONNX exports and Microsoft Foundry Local
  • Multimodal variants for text, image, and visual reasoning

Typical Use Cases

Edge AI: Deploying language models on resource-constrained devices such as IoT gateways or local servers.

Mobile Applications: AI features in apps without a cloud roundtrip for better latency and offline capability.

Reasoning-Intensive Tasks: Mathematical problem solving, tutoring applications, and agentic workflows using the Phi-4-reasoning variants at a lower resource footprint than large foundation models.

Benefits

  • Significantly lower inference costs than large LLMs
  • Faster response times through compact size
  • Local execution possible for privacy and offline operation
  • Open weights available for customization and self-hosting

Frequently Asked Questions

How do Phi models differ from large foundation models?

Phi models are considerably smaller and more resource-efficient, but tend to be less versatile for very broad or highly complex tasks. They are particularly well suited for focused use cases such as reasoning, edge deployment, or mobile scenarios.

Can I run Phi models locally?

Yes, Phi models are available as open weights and ONNX exports and can run via Microsoft Foundry Local or your own infrastructure without a cloud connection.

What does using Phi models cost?

The model weights can be downloaded and self-hosted free of charge. If you use the models as a managed service via Microsoft Foundry, billing depends on the deployment type: per token usage or per compute hour for dedicated endpoints.

Which model variants are currently available?

The Phi family currently includes Phi-4, the reasoning variants Phi-4-reasoning, Phi-4-reasoning-plus, and Phi-4-mini-reasoning, the latency-optimized Phi-4-mini-flash-reasoning, and Phi-4-reasoning-vision for multimodal tasks. The exact catalog selection changes as new releases ship.

How do I use Phi in Azure?

Phi models are available via the Microsoft Foundry model catalog and can be deployed as a serverless endpoint billed per token or as a managed compute deployment.

Integration with innFactory

As a Microsoft Solutions Partner, innFactory supports you with Phi models: evaluation for your use case, fine-tuning, edge deployment, and integration into existing applications.

Typical Use Cases

Edge AI applications
Resource-efficient inference
Mobile AI solutions
On-device language models

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Azure (official documentation). This page does not represent an offer by Azure.

Microsoft Solutions Partner

innFactory is a Microsoft Solutions Partner. We provide expert consulting, implementation, and managed services for Azure.

Microsoft Solutions Partner Microsoft Data & AI

Similar Products from Other Clouds

Other cloud providers offer comparable services in this category. As a multi-cloud partner, we help you choose the right solution.

Google Cloud

Agent Development Kit (ADK) - Multi-Agent Framework

Agent Development Kit (ADK): Google's open-source framework to build, evaluate, and deploy single- and multi-agent …

Pricing Free / open source (Apache 2.0); …
SLA N/A (framework); SLA depends on the chosen deployment target
Compare →
Google Cloud

Agent Search (formerly Vertex AI) - AI Enterprise Search

Agent Search, formerly Vertex AI Search, provides AI-powered enterprise search with natural language and semantic …

Pricing Pay-per-use (e.g. per query and storage)
SLA SLA as published by the provider
Compare →
Google Cloud

Agent Studio - Enterprise AI Agents (ex Agent Builder)

Agent Studio (formerly Vertex AI Agent Builder) creates AI agents with RAG and grounding on enterprise data in the …

Pricing Pay-per-use
SLA SLA as published by the provider
Compare →
Google Cloud

Agent Studio (ex Vertex AI) - Generative AI Development

Agent Studio, formerly Vertex AI Studio, is Google's workspace for generative AI: prompt design, model tuning, and …

Pricing Pay-per-use
SLA SLA as published by the provider
Compare →
AWS

Amazon Augmented AI (A2I) - Human Review for ML

Amazon Augmented AI (A2I) enables human review of ML predictions. Human-in-the-loop workflows for AI quality assurance.

Pricing Pay-per-use: price per human review task
SLA SLA as published by the provider (A2I is part of Amazon SageMaker AI)
Compare →
AWS

Amazon Bedrock AgentCore - AI Agent Runtime

Amazon Bedrock AgentCore: serverless runtime and services to securely run, scale, govern and observe production AI …

Pricing Pay-per-use (consumption-based, …
SLA N/A
Compare →

74 comparable products found across other clouds.

Ready to start with Phi Open Models - Microsoft Small Language Models?

Our certified Azure experts help you with architecture, integration, and optimization.

Schedule Consultation