Fam
Data Engineer (SDE 2)
- Location
- Bengaluru
- Stipend
- 20-35 LPA
- Type
- Full-time
About this role
Batch: 2020/2021/2022/2023. About the role: - Support and drive engineering initiatives across the data platform, from technical requirement documents and design reviews through to implementation - Design and maintain data models for analytical systems, including Fact and Dimension tables, slowly changing dimensions (SCD types), cumulative tables, and One Big Table patterns - Build and optimize data pipelines that ingest, transform, and serve data at scale across batch and real-time scenarios - Partner with Product and business stakeholders to understand requirements, translate them into data solutions, and deliver measurable business value through analytics What you'll work on: - Architecting and implementing cloud-native data systems on AWS, using services like S3, EKS, and MSK with Infrastructure-as-Code practices - Writing and maintaining data processing code in Python and PySpark, complemented by strong SQL skills for both querying and optimization - Orchestrating complex workflows using Airflow or Temporal, ensuring reliability and observability across data operations - Designing and tuning Lakehouse table formats such as Iceberg, Delta, and Hudi to balance query performance with storage efficiency - Extracting and syncing high-throughput data streams from relational databases into Lakehouse systems using tools like Debezium or PeerDB - Building robust ETL and ELT frameworks capable of handling both batch and streaming workloads in PySpark and Flink - Managing and optimizing data analytics queries on engines like Trino or Clickhouse to power real-time dashboards in Metabase, Superset, and PowerBI What we're looking for: - 3–5 years of hands-on experience in Data Engineering, with proven expertise in distributed systems and cloud-native architectures - Expert proficiency in Python and PySpark; advanced SQL capabilities; familiarity with Go, Java, or Scala is valued - Demonstrated ability to reason about trade-offs when selecting storage formats and processing frameworks for different use cases - Comfortable using AI-native tools—Claude, Codex, Copilot—to accelerate code writing, testing, and documentation - Experience managing AWS data services (EMR, EKS, MSK, S3) and automating deployments via Gitlab or Github Actions - A self-starting mindset: you can lead technical direction, propose and champion new ideas, and take ownership of domain areas - Strong communication skills to interact with cross-functional teams, understand business needs, and articulate technical recommendations
How well do you fit this role?
Your résumé against this posting — what you have, what’s missing, and a short plan to close the gap.