Building automation, data engineering, and integration solutions that solve real business problems.
I specialize in turning high-velocity event streams into reliable, AI-ready data products. With hands-on experience across the full data lifecycle, I design, build, and support production-grade data platforms.
I am a Software Engineer with experience in Python development, automation, data engineering, API integrations, ETL workflows, and enterprise platform support. I specialize in building automation solutions, developer tools, data validation frameworks, and integration systems. Recently focused on open-source contributions and developing DataOps Toolkit, a suite of tools for SQL lineage analysis, schema comparison, validation, metadata profiling, and reporting.
My core focus is on configuration-driven data platforms, real-time pipelines, lakehouse architectures, and the engineering discipline that makes production systems reliable — testing, CI/CD, observability, and security.
# Platform in a nutshell data_source = "Kafka" processing = "PySpark Structured Streaming" storage = "Delta Lake" ci_cd = "GitLab CI/CD" testing = "pytest, comprehensive coverage" platform = "AI-ready, scalable, secure"
Quantifiable impact from production data engineering at Accenture.
Hands-on development and validation of configuration-driven data pipelines serving enterprise data products.
Real-world solutions with measurable business impact and production-grade reliability.
Problem: Enterprise toolkit for SQL validation, lineage analysis, schema comparison, and data quality auditing.
Outcome: 6 production modules, 23 tests, published to GitHub with B+ grade after maintainer audit.
Problem: Fixed critical bug in Dagster's AssetNode parent_keys handling affecting data pipeline reliability.
Outcome: PR #34143 submitted, regression tests added, improving reliability for thousands of users.
Problem: Build scalable real-time streaming pipelines with exactly-once delivery guarantees.
Outcome: High-throughput data processing with exactly-once delivery and Delta checkpointing.
Problem: Build multi-environment cloud infrastructure for real-time data processing.
Outcome: IaC-managed platform with exactly-once streaming and Delta Lake checkpointing.
Problem: Build intelligent document Q&A with retrieval-augmented generation.
Outcome: FastAPI + Streamlit app with vector search and configurable LLM backends.
Problem: Scale data pipeline development across enterprise business domains.
Outcome: Delivered numerous production jobs, improved deployment efficiency, maintained high availability, achieved comprehensive test coverage.
End-to-end data engineering capabilities from streaming pipelines to AI-ready platforms.
ETL/ELT pipelines, data modeling, and data quality solutions.
Real-time event streaming and exactly-once processing.
Production-grade Spark applications with optimization.
Lakehouse architecture and Delta Lake implementations.
Infrastructure as Code for data platforms.
High-performance APIs and microservices.
Retrieval-augmented generation and vector search.
End-to-end data platform design and operations.
14+ Public Pull Requests across major open-source projects. I don't just build projects — I contribute to production software used by millions.
Core technologies I use to build, test, and operate data platforms.
Industry-recognized credentials that back the skills and projects displayed here.
Accenture · Dec 2023 - Feb 2024
Elite · 73% · IIT Kharagpur · Jul–Oct 2022
Verify Credential →Deep dives into data platform engineering, streaming patterns, and career learnings.
Hands-on guide to production PySpark Structured Streaming with Kafka, Delta Lake, and cloud deployment.
Read Article →How to build a retrieval-augmented generation chatbot with document ingestion, vector search, and LLM generation.
Read Article →My study plan, resources, tips, and what actually helped me pass the Databricks Data Engineer Associate exam.
Read Article →Madanapalle Institute of Technology & Science, Madanapalle, AP
2019–2023 · 7.97 CGPA · First Class Distinction
Sri Vivekananda Junior College, Chittoor, AP
2017–2019 · 9.67 CGPA
Z.P.H.S Kanipkam, Chittoor, AP
2016–2017 · 9.3 CGPA
94.17 Percentile · 2019
Quantifiable results and engineering discipline that delivers reliable data systems.
Location: Bengaluru, India
Remote: Open to Remote Opportunities
Reach out if you're hiring or want to collaborate on data platforms.
Expected Response: Within 24 Hours