Data Pipeline definition
A data pipeline is an automated series of steps that moves data from one or more sources to a destination, transforming it along the way. Pipelines ingest data from databases, applications, files or event streams, clean and reshape it, and deliver it to warehouses, lakes, applications or machine learning systems, either in scheduled batches or continuously in real time.
How does a data pipeline work?
Every pipeline has a source, processing steps and a destination. A pipeline might read new orders from a PostgreSQL database using change data capture, enrich them with customer data from a CRM, convert currencies, validate required fields, and write the results into a warehouse table. An orchestrator schedules each step, manages dependencies between them, retries failures and alerts the team when something cannot be fixed automatically.
Pipelines run as batch jobs on a schedule, for example every hour, or as streaming jobs that process each event within seconds of it happening. The choice depends on how fresh the data must be and how much complexity and cost the team can support, since streaming systems are typically harder to build and operate. Start with batch unless a use case genuinely needs fresher data.
Key components of a data pipeline
Whatever the tools, production pipelines share a common set of building blocks. Leaving out monitoring or quality checks is the most common reason pipelines fail silently, producing dashboards that look fine while showing incomplete or incorrect numbers. Each component below should be owned, documented and covered by alerts. Review alert noise regularly so real failures are not ignored.
- Ingestion: connectors, APIs, change data capture or event streams.
- Processing: SQL, Python, Spark or stream processors like Apache Flink.
- Storage: warehouse, lake, lakehouse or operational database.
- Orchestration: Airflow, Dagster, Prefect or cloud workflow services.
- Data quality: tests on volume, schema, nulls and business rules.
- Monitoring and lineage: alerts, run history and dependency tracking.
Data pipeline vs ETL
ETL is one type of data pipeline, specifically one that extracts, transforms and loads data into a target such as a warehouse. Data pipeline is the broader term. It includes ELT pipelines, streaming pipelines that feed real-time dashboards, reverse ETL that pushes warehouse data into business tools, and machine learning pipelines that prepare features and retrain models on a schedule. The common thread is automation: data moves without anyone copying files by hand.
Example of a data pipeline
A delivery app streams driver location updates through Apache Kafka. A Flink job calculates estimated arrival times and pushes them to the customer app in seconds. In parallel, a batch pipeline copies completed orders nightly into BigQuery, where dbt models compute delivery performance by city and restaurant for operations dashboards. A separate weekly pipeline prepares training data for the model that predicts preparation times, so the estimates improve as more orders complete.
Best practices for reliable pipelines
Design pipelines to be idempotent, so rerunning a step does not duplicate data, and incremental, so they process only new or changed records. Validate data at boundaries and quarantine bad records instead of failing silently or loading them. Keep pipeline code in version control with tests and CI, define service levels for freshness, and alert on missed deadlines as well as failures. Nexzem builds batch and streaming pipelines with these practices for analytics, operations and AI workloads.