What Is the Role?
We need a data engineer who can build and operate production data pipelines on AWS. The primary focus is batch processing with S3, AWS Glue, and Athena for analytics and AI use cases.
This is a hands-on role where you own the data layer end-to-end, from ingestion and transformation to quality, security, observability, and documentation.
Key Responsibilities
Pipeline Development:
- Design and build batch ETL and ELT pipelines using AWS Glue and Amazon S3
- Ingest data from relational databases, APIs, object storage, and flat files
- Implement schema evolution, partitioning, and storage optimization using columnar file formats such as Parquet and ORC
- Orchestrate pipelines using AWS Glue workflows, Step Functions, or Amazon MWAA with Apache Airflow
Data Lake & Storage:
- Design and maintain an S3-based data lake with clear raw, cleaned, and curated layers
- Optimize S3 layouts for query performance and cost through partitioning, compaction, and lifecycle policies
- Manage schemas, metadata, and data discovery with the AWS Glue Data Catalog
Query & Analytics:
- Write and optimize Athena queries for performance and cost
- Build documented tables and views that analytics and BI teams can use independently
- Support analytical data models, including star schemas, dimensional models, and fit-for-purpose wide tables
Quality & Operations:
- Automate data quality checks and validation at each pipeline stage
- Monitor pipeline failures, data anomalies, and operational metrics with CloudWatch
- Enforce access controls, encryption, and governance using IAM, KMS, and Lake Formation
- Define and deploy data infrastructure using Terraform, AWS CDK, or CloudFormation
- Document data flows, schemas, ownership, and pipeline dependencies
Required Skills
AWS Data Services (Hands-on):
- Amazon S3: data lake storage, lifecycle policies, access control, and layout optimization
- AWS Glue: production PySpark jobs, crawlers, Data Catalog, and job bookmarks
- Amazon Athena: writing and optimizing analytical queries over data in S3
- AWS Glue workflows, Step Functions, or Amazon MWAA: pipeline orchestration
- Amazon CloudWatch: monitoring, logging, dashboards, and alerting
- IAM, KMS, and Lake Formation: access management, encryption, and data governance
Data Engineering Fundamentals:
- 2+ years building and operating data pipelines in production
- Strong SQL skills, including complex joins, window functions, CTEs, and query optimization
- Strong Python skills and practical experience building transformations with PySpark
- Experience with columnar file formats such as Parquet or ORC and effective partitioning strategies
- Understanding of layered data lake patterns such as bronze/silver/gold or raw/cleaned/curated
- Experience with star schemas, dimensional models, and analytical table design
- Automated data quality testing using AWS Glue Data Quality, Deequ, Great Expectations, or similar tools
General:
- Git and CI/CD for data pipeline code using tools such as GitHub Actions or CodePipeline
- Infrastructure as code using Terraform, AWS CDK, or CloudFormation
- Clear communication and the ability to translate business data needs into technical designs
Preferred Skills
- Experience with table formats such as Apache Iceberg, Apache Hudi, or Delta Lake
- Streaming ingestion with Amazon Kinesis or Apache Kafka, and CDC with AWS DMS or Debezium
- Familiarity with Amazon Redshift or Amazon EMR
- Experience supporting ML pipelines, feature engineering, or feature stores
- Scala experience or advanced Spark performance tuning
- Experience at a consulting or product engineering firm
Personal Qualities
- You care about data quality; bad downstream data bothers you
- You are a methodical debugger who can trace a pipeline failure from alert to root cause
- You consider performance and cost from the beginning
- You document data flows and schemas without being asked
- You work comfortably across analytics, ML, product, and engineering teams
What We Offer
- Opportunity to work on GenAI and cloud-first projects for diverse clients
- Collaborative engineering culture with mentoring and career growth
- Location-adjusted compensation and benefits
- Flexible work arrangements