Must Have Skills
- Python
- PySpark
- Apache Spark
- AWS Redshift
- AWS Glue
- AWS EMR
- AWS S3
- AWS Lambda
- SQL Development & Query Optimization
- Data Warehousing (Redshift or Hive)
- ETL/ELT Pipeline Development
- Airflow (or similar scheduling tools)
- Hadoop Ecosystem
- Data Pipeline Development (Batch & Near Real-Time)
- Data Security & Data Protection
- OLTP and OLAP Database Concepts
Good to Have Skills
- Data Lakes
- Data Modeling
- Netezza
- Informatica
- DynamoDB
- MongoDB
- AWS Athena
- AWS Step Functions
- Investment Banking Domain
- NoSQL Databases
Responsibilities
- Design, develop, and maintain scalable data pipelines using AWS services and PySpark.
- Create and maintain optimal data pipeline architecture for efficient data processing.
- Build data ingestion, transformation, and ETL workflows.
- Develop batch and near real-time data processing solutions.
- Work with large and complex datasets to meet business requirements.
- Manage and optimize data warehouses using AWS Redshift or Hive.
- Ensure data quality, integrity, security, and governance standards.
- Perform SQL tuning and performance optimization.
- Automate data workflows and improve platform scalability.
- Collaborate with Product, Data, Engineering, and Business teams for solution delivery.
- Create technical documentation and support operational readiness.
