Professional Experience
Data Engineer, Management & Governance Analyst
Accenture · Bengaluru, India
Feb 2024 – Present (Associate → Analyst, effective March 2026)
- Developed and maintained numerous production ETL jobs across multiple business domains through Python, JSON configuration, and DDL development.
- Led change requests across data products and completed validations by comparing design documents, source messages, and data across multiple environments.
- Worked in cross-functional data engineering teams covering multiple business domains; owned data pipeline domains while supporting live monitoring.
- Developed configuration-driven pipeline definitions and reusable Python utilities, improving deployment efficiency.
- Maintained high pipeline availability through modular PySpark pipelines with error handling, retry logic, schema validation, and data quality checks.
- Achieved comprehensive test coverage across pytest suites with mocked components and integration patterns.
- Optimized data processing performance via partitioning, caching, and query tuning.
- Owned end-to-end data quality and platform validation for sprint releases across source-to-target data layers.
- Used Jira dashboards to track sprint tasks, incidents, and maintenance activities for data products.
- Developed data products from design documents and built downstream datasets from source schemas, applying transformation queries when multiple source data products feed a single dataset; maintained per-schema exception tables to capture invalid records with target table reference, error log, and timestamp.
Selected Projects
DataOps Toolkit
- Built an enterprise-grade CLI toolkit for SQL validation, lineage analysis, schema comparison, and data quality auditing.
- Implemented multi-dialect SQL parsing, column-level lineage analysis, Hive partition auditing, and metadata profiling.
- Published to GitHub with 6 modules, 23 tests, and B+ grade after maintainer audit.
Dagster Open Source Contribution
- Fixed critical bug in Dagster's AssetNode parent_keys handling, improving data pipeline reliability for thousands of users.
- Added regression tests and scoped warning suppression for async test execution; PR #34143 submitted to Dagster repository.
RAG Document QA Chatbot
- Built dense vector retrieval with ChromaDB, FastAPI backend, response caching, and modular LLM interfaces (OpenAI/local LLMs).
Kafka → PySpark → Delta Pipeline
- Ingested JSON events from Kafka, enforced schemas, and wrote exactly-once to Delta Lake using checkpointing and idempotent writes.
- Built high-throughput streaming pipeline with comprehensive pytest coverage using an in-memory Spark fixture and continuous integration.