Skip to main content
Cloud / STACKIT / Products / STACKIT Intake - Data Ingestion for the Data Lakehouse

STACKIT Intake - Data Ingestion for the Data Lakehouse

STACKIT Intake: managed data ingestion via the Kafka protocol directly into Apache Iceberg tables for the STACKIT Data Lakehouse.

Data & AI
Pricing Model Capacity-based; capacity is configured via max. message size and max. messages/hour (product capped at 20 GiB/h); see STACKIT pricing calculator for details
Availability Availability as documented by STACKIT
Data Sovereignty GDPR-compliant, operated on sovereign STACKIT infrastructure
Reliability SLA as published by the provider SLA

What is STACKIT Intake?

STACKIT Intake is a fully managed data ingestion service that accepts high-volume data streams via the Apache Kafka protocol and writes them directly to Apache Iceberg tables. It forms the ingestion layer of the STACKIT Data Lakehouse, writing to tables managed via the Iceberg REST catalog of STACKIT Dremio (version 25 or 26). Intake is designed as a temporary buffer, not a permanent data store, and has been in Public Preview since December 2025 according to STACKIT release notes.

Core Features

  • Kafka protocol compatibility: Use existing Kafka client libraries and data producers such as Debezium without modification
  • Automatic schema management: Data types are inferred from JSON payloads, with automatic schema evolution of the target Iceberg table
  • Buffering with delivery guarantee: Messages are buffered for up to 24 hours until reliably written to Dremio
  • Dead letter queue: Automatic handling of malformed or unprocessable messages
  • Scalable throughput: Configurable via max. message size and max. messages/hour, with the product of both values capped at 20 GiB/h
  • Management via portal, CLI, and Terraform

Typical Use Cases

Data lakehouse ingestion: Intake serves as the ingestion layer: data flows via the Kafka protocol directly into Iceberg tables that can be analyzed in STACKIT Dremio.

Change data capture (CDC): Changes from operational databases are captured using tools like Debezium and brought into the data lakehouse in real time.

IoT and telemetry: Device data, sensor readings, and application logs are captured at high volume in real time and made available for downstream analytics.

Benefits

  • GDPR-compliant, operated on sovereign STACKIT infrastructure
  • Kafka protocol compatibility avoids lock-in for data producers
  • Automatic schema management reduces manual maintenance effort
  • Seamless integration with STACKIT Dremio for analysis directly on incoming data

Integration with innFactory

As an official STACKIT partner, innFactory supports you in designing real-time data pipelines: from connecting existing Kafka producers to schema design and integration with STACKIT Dremio for downstream analytics.

Typical Use Cases

Real-time data ingestion for the STACKIT Data Lakehouse (STACKIT Dremio)
Change data capture (CDC), e.g. via Debezium
IoT and telemetry data streams
Log and event aggregation

Technical Specifications

Buffering Temporary buffer, messages retained for up to 24 hours
Data format JSON, with automatic schema inference and evolution
Dremio compatibility Requires STACKIT Dremio 25 or 26 with administrative access for the integration
Protocol Apache Kafka protocol (compatible with existing Kafka clients)
Target Writes to Apache Iceberg tables via the Dremio Iceberg REST catalog
Throughput Configurable via max. message size (1-1024 KiB) and max. messages per hour; the product of both values is capped at 20 GiB/h

Frequently Asked Questions

What is STACKIT Intake?

A managed data ingestion service that accepts data streams via the Apache Kafka protocol and writes them directly to Apache Iceberg tables in the Dremio Iceberg REST catalog. It serves as a bridge between data producers and the STACKIT Data Lakehouse.

Is Intake a permanent data store?

No. Intake is designed as a temporary buffer, holding messages for up to 24 hours until they are reliably written to the target Iceberg table. They are then removed from the buffer.

What data formats are supported?

Currently, STACKIT Intake processes JSON data only. Data types are automatically inferred from JSON payloads, and the schema of the target Iceberg table is automatically managed and evolved as new fields appear.

What is the maximum throughput?

An Intake Runner's capacity is configured via two parameters: maximum message size (1 to 1024 KiB) and maximum messages per hour. The product of both values may not exceed 20 GiB/h.

Do I need to replace my existing Kafka clients?

No. Since STACKIT Intake supports the Apache Kafka protocol, existing Kafka client libraries and common data producers such as Debezium can be reused without modification.

What status does STACKIT Intake have and which Dremio version is required?

According to the release notes, STACKIT Intake is in Public Preview (since December 2025). Integration requires STACKIT Dremio 25 or 26 with administrative access.

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of STACKIT (official documentation). This page does not represent an offer by STACKIT.

STACKIT Partner

innFactory is an official STACKIT Partner. We provide consulting, implementation, and managed services for the sovereign cloud.

STACKIT Official Partner

Similar Products from Other Clouds

Other cloud providers offer comparable services in this category. As a multi-cloud partner, we help you choose the right solution.

Google Cloud

Agent Platform Workbench (formerly Vertex AI Workbench)

Agent Platform Workbench provides managed JupyterLab environments on the Gemini Enterprise Agent Platform with BigQuery, …

Pricing Usage-based via the Gemini Enterprise …
SLA SLA as published by the provider
Compare →
Google Cloud

Cloud Talent Solution - Job Search with Machine Learning

Cloud Talent Solution brings machine learning to the job search experience and supports recruiting platforms with job …

Pricing Pricing as of January 5, 2021: charged …
SLA As published by the provider / see official documentation
Compare →
Google Cloud

Deep Learning VM Images - Preconfigured ML VM Images

Deep Learning VM Images are virtual machine images optimized for data science and machine learning with frameworks such …

Pricing The images themselves are free to use …
SLA SLA as published by the provider
Compare →
Google Cloud

Enterprise Knowledge Graph - Entity Reconciliation for BigQuery

Enterprise Knowledge Graph reconciles siloed data through an entity reconciliation API and adds lookups via the Google …

Pricing Google does not publish a dedicated …
SLA Pre-GA products are available "as is" and might have limited support per the documentation; As published by the provider / see official documentation
Compare →
Google Cloud

Genkit - Open Source Framework for AI Applications

Genkit is Google's open-source framework for building full-stack, AI-powered and agentic applications with TypeScript, …

Pricing Open-source framework; costs arise from …
SLA No SLA applies to the framework itself; SLAs apply to the services used as published by their providers
Compare →
Google Cloud

Model Garden - Model Catalog of the Agent Platform

Model Garden is the model library of the Gemini Enterprise Agent Platform for discovering, testing, customizing, and …

Pricing For open source models you are charged …
SLA SLA as published by the provider
Compare →

128 comparable products found across other clouds.

Ready to start with STACKIT Intake - Data Ingestion for the Data Lakehouse?

Our certified STACKIT experts help you with architecture, integration, and optimization.

Schedule Consultation