Open to Senior AWS Data Engineering Roles

Sonu Ansari.

Senior AWS Data Engineer

Redshift · Kinesis · Glue · Lambda · Step Functions · S3 · DynamoDB · PySpark · Python

4+ years building production-grade data platforms on AWS — serverless ETL pipelines, real-time Kinesis streaming at 10K events/sec, Redshift data warehouses, and fault-tolerant orchestration with Step Functions. Delivered 80% cost reduction, <30s end-to-end streaming latency, and zero missed SLAs across healthcare and IoT fleets of 1,000+ vehicles.

4+
Years on AWS Data
80%
Pipeline Cost Cut
1K+
Vehicles Streamed
<30s
Kinesis Latency
01.

Roles I Target & Match

Senior AWS Data Engineer Strong match

GlueLambdaS3Step FunctionsPySpark

Senior AWS Data Engineer — Redshift Strong match

RedshiftSpectrumCOPY/UNLOADdist & sort keysSQL tuning

AWS Data Engineer — Kinesis Strong match

Kinesis Data StreamsLambda consumersshard scalingIoT Core

AWS Data Platform Engineer Strong match

S3 data lakeBronze/Silver/GoldIAMDelta LakeTerraform

Real-Time Data Engineer (AWS) Strong match

<30s latencyKinesisLambdaDynamoDBTimestream

Streaming Data Engineer — Kinesis Strong match

10K events/secKafkaKinesisevent-drivenzero data loss

Senior Data Engineer — Redshift + Python Strong match

PythonPySparkRedshiftSQLpandas

AWS Data Engineer — Step Functions Strong match

orchestrationretry/catchparallel statesDLQAirflow

AWS Data Engineer — S3 · Redshift · Kinesis Strong match

full AWS data stacklake → warehousestreaming + batch

Cloud Data Engineer (AWS) Strong match

AWS certifiedserverlessCI/CDDockerLinux
02.

About Me

I'm a Senior AWS Data Engineer with 4+ years of experience designing and operating production data platforms on Amazon Web Services. My core work spans serverless ETL pipelines (AWS Glue, Lambda, Step Functions), real-time streaming with Amazon Kinesis, Redshift data warehousing, S3-based data lakes, and DynamoDB/Timestream — turning high-volume batch and streaming data into reliable, decision-ready products.

In healthcare, I run concurrent production ETL workloads delivering audit-ready data with zero missed SLAs, and cut pipeline runtime and infrastructure cost by 80% ($500/month) through PySpark profiling and optimization. In IoT, I built an end-to-end telemetry platform on AWS IoT Core → Kinesis → Lambda tracking 1,000+ vehicles at 10K events/sec with <30-second latency and zero data loss — including geofencing, alerting, and Google Maps road-matching.

I care about what happens beyond making a pipeline run: data quality gates, observability, fault-tolerant orchestration, cost optimization, and the experience of downstream consumers. Whether it's tuning Redshift distribution keys, scaling Kinesis shards, rewriting a Spark job, or designing a Bronze/Silver/Gold lakehouse, I like understanding the problem deeply and shipping clean, working solutions.

AWS · Redshift · Kinesis · Glue · Lambda · Step Functions · S3 · DynamoDB · Timestream · IoT Core · PySpark · Python · SQL · Kafka · Airflow · Databricks · Delta Lake · Terraform · Docker · CI/CD
03.

Experience

04.

Skills — AWS Data Engineering Focus

⭐ AWS Data Engineering (Core)
Amazon Redshift Amazon Kinesis AWS Glue AWS Lambda AWS Step Functions Amazon S3 Amazon DynamoDB Amazon Timestream AWS IoT Core Amazon EC2
⭐ Streaming & Real-Time
Kinesis Data Streams Apache Kafka Event-driven architectures Real-time alerting Geofencing IoT telemetry
Processing & ETL
PySpark Databricks Delta Lake Apache Airflow dbt Pandas NumPy
Databases & Warehousing
Redshift PostgreSQL MySQL MongoDB Redis Delta Lake
Languages
Python SQL Bash JavaScript TypeScript Java
DevOps, Quality & Delivery
Terraform Docker Git CI/CD Robot Framework Data-quality testing Linux Jira Confluence Agile / Scrum
05.

Personal Projects

📷
Edge Face-Recognition Surveillance (Raspberry Pi)
Home surveillance prototype on Raspberry Pi + OpenCV: real-time face recognition, activity recording, and owner alerts for unknown persons — edge computing experience that translates directly to IoT data pipelines.
  • Real-time detection with OpenCV on edge hardware
  • Known-face enrollment & recognition history
  • Automatic recording + remote phone alerts
EdgeProcessing
24/7Monitoring
Real-timeAlerts
PythonOpenCVRaspberry PiComputer Vision
🗺
Autonomous Navigation & Pathfinding
Robotics project for autonomous movement in constrained environments — shortest-path routing, grid-based planning, and sensor-driven control. National-level finalist at IIT Bombay's robotics competition.
  • A* / Dijkstra / BFS routing under obstacles
  • Sensor feedback control loops
  • IIT Bombay national competition finalist
PythonAlgorithmsRoboticsEmbedded
06.

Live Demo — AWS ETL Pipeline Architecture

From user activity to monthly reports — Kinesis → DynamoDB → Lambda → Glue → Step Functions → S3 Gold → Email

An interactive walkthrough of the serverless ETL architecture I build with: subscription filtering before expensive processing, high-throughput ingest, parsing/validation, Spark-based feature engineering, orchestrated monthly workflows, and medallion-style S3 layers feeding personalized email delivery.

Ready, filtering active subscribers
👥 User pool100,000registered customers
✓ Policy filter18,420active subscribers
📈 Monthly reports0eligible users processed
✉ Delivery0%reports emailed
👥
User Pool
100K users
→
🔐
Subscription Filter
active policy only
→
📡
Kinesis
10K events/sec
→
🗃️
DynamoDB
raw activity
→
λ
Lambda
read + parse
→
⚙
Glue / Spark
metrics + features
→
🔀
Step Functions
orchestration
→
🥉
Bronze → Silver → Gold
S3 medallion layers
→
🧠
Insights Engine
observations + badges
→
✉
Email Delivery
monthly report
0Events read
0Users validated
0Gold report rows
—Pipeline status
Why this is production-grade ETL: subscription eligibility is applied before expensive processing; Kinesis handles high-throughput ingest; Lambda parses and validates; Glue/Spark builds reusable metrics; Step Functions coordinates the fault-tolerant monthly workflow; and S3 Bronze/Silver/Gold layers separate raw, trusted, and business-ready data — with the Gold layer feeding insights and personalized email delivery.
07.

Certifications

☁️

AWS Certified Cloud Practitioner

Amazon Web Services — cloud fundamentals, architecture, security, and cost best practices

📊

Databricks Certified Associate Data Engineer

Databricks — Spark, Delta Lake, and production data pipeline development

08.

Education

🎓

Bachelor's Degree — Computer Science & Engineering

Goa University

09.

Recognition

🏆

National-Level Robotics Competition, IIT Bombay

Selected among hundreds of teams nationwide to compete in IIT Bombay's engineering challenge. Designed shortest-path routing logic for an autonomous construction-material transport robot.

Let's Build Your Data Platform

Open to Senior AWS Data Engineer, Redshift, Kinesis, and streaming roles. If you're hiring at Amazon, R Systems, ApTask, TEKsystems, Cloud Kinetics, Cognizant, or Quest Global — I'd love to talk.