Skip to main content
Cloud / Azure / Products / Foundry Observability - AI Monitoring & Evaluation

Foundry Observability - AI Monitoring & Evaluation

Foundry Observability provides tracing, evaluation, and monitoring for AI applications and agents on Microsoft Foundry.

ai-machine-learning
Pricing Model Consumption-based billing, e.g. for agent playground evaluations and risk & safety evaluations
Availability Global, availability of individual evaluators varies by region
Data Sovereignty EU regions available
Reliability SLA as published by the provider (see the official Microsoft Foundry SLA page) SLA

What is Foundry Observability?

Foundry Observability is the observability component of Microsoft Foundry for AI applications and agents. The service combines three core capabilities: Evaluation measures the quality, safety, and reliability of AI responses using built-in evaluators (including coherence, groundedness, relevance, safety and fairness metrics, and agent-specific metrics such as tool call accuracy). Monitoring delivers real-time dashboards on operational metrics, token consumption, latency, error rates, and quality scores through integration with Azure Monitor Application Insights. Tracing, built on OpenTelemetry, captures the full execution path of a request across model calls, tool usage, and agent decisions.

Observability accompanies the entire AI application lifecycle: from model selection, through pre-production evaluation using your own test datasets, to continuous monitoring in production, including scheduled evaluations for drift detection and an AI red teaming agent for automated security testing.

Core Features

  • OpenTelemetry-based distributed tracing for LLM and tool calls, with support for frameworks such as LangChain, LangGraph, the OpenAI Agents SDK, and Microsoft Agent Framework
  • Built-in evaluators for quality (coherence, fluency), RAG-specific metrics (groundedness, relevance), safety (hate/unfairness, violence, protected materials), and agent behavior (tool call accuracy, task completion)
  • Custom evaluators for domain-specific requirements
  • Production monitoring with real-time dashboards via Azure Monitor Application Insights
  • Scheduled evaluations and an AI red teaming agent (based on the PyRIT framework) for security testing before and after deployment

Typical Use Cases

Engineering teams monitor LLM and agent applications in production, track latency spikes and error rates, and use traces to identify slow or failing tool calls.

Responsible AI and compliance teams use built-in safety and fairness evaluators along with the AI red teaming agent to test applications for vulnerabilities before rollout.

ML and platform teams monitor response quality over time using scheduled evaluations to detect drift, for example when groundedness scores drop and a model starts to hallucinate.

FinOps teams use operational metrics on token consumption to track costs per use case or feature.

Benefits

  • End-to-end visibility across the entire AI application lifecycle, from model selection to production operation
  • Standardized tracing based on open OpenTelemetry standards, compatible with common agent frameworks
  • Prebuilt evaluators reduce effort for quality and safety assessment
  • Seamless integration with Azure Monitor Application Insights for alerting and dashboards
  • Supports compliance and Responsible AI requirements through red teaming and safety evaluators

Integration with innFactory

As a Microsoft Solutions Partner, innFactory supports you with Foundry Observability: monitoring setup, selecting and configuring the right evaluators, dashboard design, and alert configuration for your AI and agent applications.

Contact us for a non-binding consultation on Foundry Observability and Microsoft Azure.

Typical Use Cases

Monitoring LLM and agent applications in production
Quality and safety evaluation before and after deployment
Distributed tracing for model calls and tool usage
Detecting quality drift and hallucinations over time

Frequently Asked Questions

What is Foundry Observability?

Foundry Observability is the observability component of Microsoft Foundry and comprises three core capabilities: evaluation of response quality and safety, production monitoring with real-time dashboards, and distributed tracing of model and tool calls.

What is captured and logged?

Captured data includes traces of LLM calls, tool usage, and agent decisions, operational metrics such as latency, error rates, and token consumption, as well as evaluation results on quality and safety. Tracing is based on OpenTelemetry.

How do I integrate Observability into my application?

Tracing is captured via OpenTelemetry standards and supports common frameworks such as LangChain, LangGraph, the OpenAI Agents SDK, and Microsoft Agent Framework. Monitoring and evaluation results are integrated with Azure Monitor Application Insights.

Can I set alerts for quality issues?

Yes, Azure Monitor lets you configure notifications when outputs fall below defined quality thresholds or produce harmful content.

What does Foundry Observability cost?

Billing is consumption-based, for example for risk and safety evaluations and evaluations run in the agent playground. Playground evaluations are enabled by default and billed as part of consumption, but can be disabled.

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Azure (official documentation). This page does not represent an offer by Azure.

Microsoft Solutions Partner

innFactory is a Microsoft Solutions Partner. We provide expert consulting, implementation, and managed services for Azure.

Microsoft Solutions Partner Microsoft Data & AI

Similar Products from Other Clouds

Other cloud providers offer comparable services in this category. As a multi-cloud partner, we help you choose the right solution.

Google Cloud

Agent Development Kit (ADK) - Multi-Agent Framework

Agent Development Kit (ADK): Google's open-source framework to build, evaluate, and deploy single- and multi-agent …

Pricing Free / open source (Apache 2.0); …
SLA N/A (framework); SLA depends on the chosen deployment target
Compare →
Google Cloud

Agent Search (formerly Vertex AI) - AI Enterprise Search

Agent Search, formerly Vertex AI Search, provides AI-powered enterprise search with natural language and semantic …

Pricing Pay-per-use (e.g. per query and storage)
SLA SLA as published by the provider
Compare →
Google Cloud

Agent Studio - Enterprise AI Agents (ex Agent Builder)

Agent Studio (formerly Vertex AI Agent Builder) creates AI agents with RAG and grounding on enterprise data in the …

Pricing Pay-per-use
SLA SLA as published by the provider
Compare →
Google Cloud

Agent Studio (ex Vertex AI) - Generative AI Development

Agent Studio, formerly Vertex AI Studio, is Google's workspace for generative AI: prompt design, model tuning, and …

Pricing Pay-per-use
SLA SLA as published by the provider
Compare →
AWS

Amazon Augmented AI (A2I) - Human Review for ML

Amazon Augmented AI (A2I) enables human review of ML predictions. Human-in-the-loop workflows for AI quality assurance.

Pricing Pay-per-use: price per human review task
SLA SLA as published by the provider (A2I is part of Amazon SageMaker AI)
Compare →
AWS

Amazon Bedrock AgentCore - AI Agent Runtime

Amazon Bedrock AgentCore: serverless runtime and services to securely run, scale, govern and observe production AI …

Pricing Pay-per-use (consumption-based, …
SLA N/A
Compare →

74 comparable products found across other clouds.

Ready to start with Foundry Observability - AI Monitoring & Evaluation?

Our certified Azure experts help you with architecture, integration, and optimization.

Schedule Consultation