Quick Start¶
Get up and running with streamt in 5 minutes. By the end of this guide, you'll have a working streaming pipeline.
Prerequisites¶
- streamt installed
- Docker running (for local Kafka)
1. Start Local Infrastructure¶
# Start Kafka, Flink, and supporting services
docker compose up -d
# Wait for services to be healthy
docker compose ps
2. Create Your Project¶
This creates stream_project.yml, sources/, models/, and tests/ directories.
Already have Kafka topics?
Use streamt init --discover --kafka localhost:9092 to auto-generate sources from existing topics. Add --schema-registry http://localhost:8081 to extract column definitions from Avro schemas.
Edit the generated configuration file:
project:
name: my-first-pipeline
version: "1.0.0"
description: My first streaming pipeline with streamt
runtime:
kafka:
bootstrap_servers: localhost:9092
flink:
default: local
clusters:
local:
type: rest
rest_url: http://localhost:8082
sql_gateway_url: http://localhost:8084
Using Confluent Cloud? Add authentication to your runtime config:
Store credentials inruntime: kafka: bootstrap_servers: pkc-abc12.us-east-1.aws.confluent.cloud:9092 security_protocol: SASL_SSL sasl_mechanism: PLAIN sasl_username: my-api-key sasl_password: my-api-secret schema_registry: url: https://psrc-xyz99.us-east-1.aws.confluent.cloud username: sr-api-key password: sr-api-secret.env(gitignored) and reference with${VAR}syntax.
3. Define a Source¶
Create a source representing incoming data:
sources:
- name: raw_events
topic: events.raw.v1
description: Raw user events from the web application
owner: platform-team
columns:
- name: event_id
description: Unique event identifier
- name: user_id
description: User who triggered the event
- name: event_type
description: Type of event (click, view, purchase)
- name: timestamp
description: When the event occurred
4. Create Your First Model¶
Create a model that transforms the raw events:
models:
- name: events_clean
description: Cleaned and validated events
sql: |
SELECT
event_id,
user_id,
event_type,
`timestamp`
FROM {{ source("raw_events") }}
WHERE event_id IS NOT NULL
AND user_id IS NOT NULL
# Optional: customize topic settings
topic:
name: events.clean.v1
partitions: 6
The model is automatically materialized as a topic since it's a simple SELECT statement.
5. Validate Your Project¶
Check that everything is configured correctly:
You should see:
6. View the Lineage¶
See how data flows through your pipeline:
Output:
7. Plan the Deployment¶
See what will be created:
Output:
Plan: 1 to create, 0 to update, 0 to delete
Topics:
+ events.clean.v1 (6 partitions, replication: 1)
8. Deploy!¶
Apply your pipeline to the infrastructure:
Output:
Applying changes...
Topics:
+ events.clean.v1 ............... created
Applied: 1 created, 0 updated, 0 unchanged
9. Verify in Conduktor¶
Open http://localhost:8080 and log in with:
- Email: admin@localhost
- Password: Admin123!
You should see the events.clean.v1 topic in the Topics view.
10. Add a Test¶
Create a test to validate data quality:
tests:
- name: events_schema_validation
model: events_clean
type: schema
assertions:
- not_null:
columns: [event_id, user_id, event_type]
- accepted_values:
column: event_type
values: [click, view, purchase, signup]
Run the test:
Project Structure¶
Your project should now look like this:
my-streaming-project/
├── stream_project.yml # Main configuration
├── sources/
│ └── events.yml # Source definitions
├── models/
│ └── events_clean.yml # Model definitions
└── tests/
└── events_test.yml # Test definitions
Single-File vs Multi-File¶
streamt supports both layouts:
| Layout | When to use |
|---|---|
Single-file (stream_project.yml with everything) |
Small projects, quick prototyping, < 5 models |
Multi-file (separate sources/, models/, tests/ dirs) |
Team projects, > 5 models, better git diffs |
Both are equivalent — streamt auto-discovers YAML files in subdirectories. You can also mix: keep sources inline in stream_project.yml and split models into models/. Subdirectory nesting works too (models/payments/orders.yml).
Bonus: Inspect Your Pipeline¶
Use list and show to explore what you've built:
# List all models
streamt list models
# Show details of a specific model
streamt show model events_clean
# Get JSON output (for scripting or LLM agents)
streamt -o json list sources
streamt -o json show model events_clean
What's Next?¶
Congratulations! You've created your first streaming pipeline with streamt.
- Build a complete pipeline — Add stateful processing with Flink
- Learn about concepts — Understand sources, models, tests
- Explore materializations — Topics, Flink jobs, sinks
- See examples — Real-world pipeline examples
- CI/CD Integration — GitHub Actions, validation in PRs
- CLI Reference — All commands and options