Role & responsibilities
- 5+ years of experience in Data Engineering and production-scale data platforms.
- Design and manage Apache Airflow pipelines for high-volume sensor, satellite, and third-party data ingestion.
- Build and optimize Apache Spark workloads for batch processing, geospatial analytics, and large-scale aggregations.
- Develop and maintain PostgreSQL, TimescaleDB, PostGIS, and DuckDB-based data solutions.
- Implement efficient data ingestion, transformation, and bulk-loading pipelines.
- Containerize and deploy applications using Docker and Docker Compose, including air-gapped environments.
- Collaborate with customer teams to translate business requirements into data products and analytical solutions.
- Ensure data quality, lineage, observability, performance, and SLA compliance.
- Troubleshoot production issues and optimize pipeline, database, and query performance.
- Create and maintain technical documentation, data models, and operational runbooks.
- Demonstrate strong proficiency in Python, SQL, Airflow, Spark, and PostgreSQL.
- Experience with geospatial and time-series data, S3-compatible storage, and secure customer environments.
- Strong communication, stakeholder management, and problem-solving skills.
