The Video Intelligence API automatically analyzes videos and extracts structured metadata for media workflows, content discovery, and moderation.
What is the Video Intelligence API?
The Video Intelligence API is Google’s pre-trained service for automatic video analysis, often also referred to as “Video AI”. The service detects objects, scenes, actions, text, and explicit content in stored or live-streamed videos, delivering structured metadata with timestamps at the video, segment, shot, and frame level.
Unlike manual video categorization, the API analyzes large volumes of content in a short time. It identifies a wide range of objects and concepts, providing timestamps and confidence scores for each detection.
For specialized requirements, AutoML Video is available. It lets you train custom models for your own object classes or classifications, such as specific product categories or industry-specific content. The API’s earlier celebrity recognition feature is considered deprecated and should be avoided for new projects.
Core Features
- Label Detection: Automatic detection of objects, places, and activities in video with timestamps.
- Shot Change Detection: Detection of camera and cut changes for automatic segmentation.
- Explicit content detection: Automatic screening for potentially inappropriate content.
- Speech Transcription: Transcription of spoken content for subtitles and search.
- Text detection (OCR) and logo detection: Recognition of text and brand/company logos in the frame.
- Object tracking and person detection: Tracking objects and people across the video timeline.
Common Use Cases
Content Moderation for Platforms
A video platform uses explicit content detection to automatically review uploaded videos. Problematic content is flagged before publication. The moderation team focuses on flagged segments instead of manual full review.
Video Cataloging and Search
A media company indexes its video archive. Label Detection recognizes objects, scenes, and activities. Editors find relevant clips by searching for terms like “office interview” or “outdoor sports scene”.
Automatic Subtitling
An e-learning provider uses Speech Transcription for automatic subtitles. The API transcribes spoken content with timestamps. This saves manual transcription effort and improves accessibility.
Logo Tracking in Broadcasts
A sponsor tracking service analyzes sports broadcasts for logo visibility. Logo Detection measures how often and how long sponsor logos appear on screen.
Shot-based Video Segmentation
A post-production company uses Shot Change Detection for automatic cut detection. The API identifies camera changes and provides the basis for further editing steps such as color grading.
Benefits
- Fast integration: Pre-trained models deliver usable results immediately without in-house ML expertise.
- Broad feature set: From object detection to transcription to logo recognition in a single API.
- Scalability: Processing large volumes of video without your own infrastructure.
- Extensibility: AutoML Video enables custom models for specialized requirements.
- Google Cloud integration: Native connection to Cloud Storage and other Google Cloud services.
Integration with innFactory
As a certified Google Cloud partner, innFactory supports you in integrating the Video Intelligence API into your media workflows: from architecture through implementation to optimization.
Contact us for a consultation.
Available Tiers & Options
Video Intelligence API
- Pre-trained models
- No ML expertise required
- Fast integration
- Limited customization
AutoML Video
- Custom model training
- Own object classes
- Higher accuracy
- Requires training data
Typical Use Cases
Technical Specifications
Frequently Asked Questions
What is the Video Intelligence API?
The Video Intelligence API (often referred to as "Video AI") automatically analyzes videos and extracts metadata. The service detects objects, scenes, text, and explicit content. AutoML Video is available for custom requirements.
What analysis features does the Video Intelligence API offer?
The API offers Label Detection (objects/actions), Shot Change Detection, explicit content detection, speech transcription, text detection (OCR), logo detection, object tracking, and person detection. Celebrity recognition is now considered deprecated and should not be used for new projects.
How does billing work?
The Video Intelligence API bills per analyzed minute, with a monthly free quota per feature. Cost per minute varies by feature and by whether you use stored-video or streaming annotation. Current prices are available on the official pricing page.
Can I train custom recognition models?
Yes, with AutoML Video you can train custom models for object detection and classification. This requires labeled training data for the desired classes.
Is the Video Intelligence API suitable for live streaming?
In addition to batch processing of stored videos, the API supports streaming annotation for live content, though with a more limited feature set and different costs compared to stored-video analysis.
Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Google Cloud (official documentation). This page does not represent an offer by Google Cloud.
