Data Platform Engineer
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.
As a Data Platform Engineer, you will own key segments of the data flow from source to trusted dataset: ingestion, change data capture, streaming, lakehouse storage, orchestration, transformation, contracts, quality, and lineage. Some data sources may involve machines, PLCs, historians, and industrial protocols. Controls experience is advantageous but not required.
This is a data-platform and distributed-systems role. Your work should enable new sources to be integrated easily, make failures straightforward to repair, and ensure datasets are reliable for use by Analytics, Data Science, Operations Research, and ML systems.
What You'll Do
- Build the data backbone for autonomous factories, transforming machine signals, quality events, work orders, and application changes into trusted data used by scheduling, ML, and operations.
- Build ingestion, CDC, and streaming capabilities for transactional data, events, telemetry, and files; explicitly manage ordering, deletes, retries, replay, idempotence, and backpressure.
- Define versioned data and event contracts with upstream teams, supported by testing and service targets for freshness, completeness, and correctness.
- Model telemetry, quality events, work orders, and operational data into datasets with explicit grain, identity, time, provenance, and history.
- Own Dagster orchestration, dbt transformation, data CI/CD, backfills, lineage, observability, and offline feature datasets for ML.
- Collaborate with Manufacturing Operations and Infrastructure to acquire data from machines, PLCs, historians, OPC-UA, MTConnect, and MQTT sources as needed.
What We're Looking For
- Experience building and operating production data infrastructure or distributed data systems, including on-call ownership and recovery efforts.
- Strong production Python and advanced SQL and data-modeling skills, including incremental processing, temporal data, and schema evolution.
- Experience with Kafka or another event-streaming platform, plus CDC or other stateful incremental pipelines.
- Experience operating Snowflake, and with a lakehouse table format such as Iceberg, Delta, or Hudi, including expertise in partitioning and compaction.
- Experience with tools such as Dagster, Airflow, Argo, or Prefect; dbt or similar transformation frameworks; and Kubernetes or infrastructure as code.
- Strong judgment regarding contracts, failure modes, and the needs of downstream analytics, ML, and operational systems.
What Will Set You Apart
- Experience running Snowflake and Iceberg together or designing a hybrid warehouse and lakehouse architecture.
- Production experience with PeerDB, Debezium, Flink, Spark Structured Streaming, Redpanda, Bufstream, or similar CDC and streaming systems.
- Proficiency with ClickHouse or another low-latency analytical database, including performance tuning and lifecycle management.
- Experience with industrial or edge data collection using OPC-UA, MTConnect, MQTT, historians, PLCs, or handling intermittently connected systems.
- Background in performance-sensitive data systems built with Go, Rust, or Scala; regulated-environment experience; or contributions to db.
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.