Open to Python, Automation, Data Engineering, and System Integration Roles
Python | Automation | Data Engineering | System Integration

Sasidhar Mopuru

Building automation, data engineering, and integration solutions that solve real business problems.

Current Focus
Apache Kafka PySpark Delta Lake Databricks Terraform FastAPI RAG Systems Open Source Contributor
PySpark Apache Kafka Delta Lake Databricks Python CI/CD
Sasidhar Mopuru
Databricks Certified
Python
Development
Automation
Engineering
Data
Engineering
System
Integration
About

Engineering Data Platforms at Scale

I specialize in turning high-velocity event streams into reliable, AI-ready data products. With hands-on experience across the full data lifecycle, I design, build, and support production-grade data platforms.

I am a Software Engineer with experience in Python development, automation, data engineering, API integrations, ETL workflows, and enterprise platform support. I specialize in building automation solutions, developer tools, data validation frameworks, and integration systems. Recently focused on open-source contributions and developing DataOps Toolkit, a suite of tools for SQL lineage analysis, schema comparison, validation, metadata profiling, and reporting.

My core focus is on configuration-driven data platforms, real-time pipelines, lakehouse architectures, and the engineering discipline that makes production systems reliable — testing, CI/CD, observability, and security.

# Platform in a nutshell
data_source = "Kafka"
processing = "PySpark Structured Streaming"
storage = "Delta Lake"
ci_cd = "GitLab CI/CD"
testing = "pytest, comprehensive coverage"

platform = "AI-ready, scalable, secure"
Experience

Key Achievements

Quantifiable impact from production data engineering at Accenture.

Open
Source
Developer
Tooling
ETL
Pipelines
API
Integration
Python
Automation
Data
Validation
Experience

Accenture — Enterprise Data Engineering

Hands-on development and validation of configuration-driven data pipelines serving enterprise data products.

Data Engineer, Management and Governance Analyst

Accenture · Feb 2024 – Present (Associate to Analyst, Feb 2026, effective March 2026) · Bengaluru, India
  • Delivered numerous production ETL jobs across multiple business domains through Python development, JSON configuration, and DDL development.
  • Led change requests across data products and completed validations by comparing design documents, source messages, and data across multiple environments.
  • Worked in cross-functional data engineering teams covering multiple business domains; owned data pipeline domains while supporting live monitoring.
  • Managed full lifecycle of data products: PySpark/Python main scripts, JSON configs, unit tests, production database tables, Delta Lake lakehouse tables (with catalog and schemas), enterprise scheduling tools, job orchestration platforms, and monitoring solutions across multiple environments.
  • Improved deployment efficiency by developing configuration-driven pipeline definitions and reusable Python utilities used by the data platform and consumed by existing CI/CD workflows.
  • Maintained high pipeline availability through modular PySpark pipelines with error handling, retry logic, schema validation, and data quality checks.
  • Maintained comprehensive test coverage across pytest suites with mocked components and integration patterns.
  • Optimized pipeline performance through partitioning, caching, and performance tuning on production pipelines.
  • Owned end-to-end data quality and platform validation for 3–4 sprint releases, validating schemas, tags, record counts, primary keys, business hash keys, and duplicate records across source-to-target data layers, and prepared SQL-based reconciliation evidence for clean production sign-off.
  • Worked with Kafka-based data pipelines and validation workflows, DDL scripts, and validation/reconciliation queries for real-time data processing, and monitored and troubleshot production pipelines.
  • Used Jira dashboards to track sprint tasks, incidents, and maintenance activities for data products.
  • Developed data products from design documents and built downstream datasets from source schemas, applying transformation queries when multiple source data products feed a single dataset; maintained per-schema exception tables to capture invalid records with target table reference, error log, and timestamp.
PySparkKafkaDelta LakeDatabricksProduction DatabasesJob OrchestrationMonitoringGitLab CI/CDJiraPython
Projects

Production-Grade Data Platforms

Real-world solutions with measurable business impact and production-grade reliability.

DataOps Toolkit ⭐

Problem: Enterprise toolkit for SQL validation, lineage analysis, schema comparison, and data quality auditing.

Outcome: 6 production modules, 23 tests, published to GitHub with B+ grade after maintainer audit.

PythonSQLsqlglotPydantic

Dagster OSS Contribution ⭐

Problem: Fixed critical bug in Dagster's AssetNode parent_keys handling affecting data pipeline reliability.

Outcome: PR #34143 submitted, regression tests added, improving reliability for thousands of users.

PythonDagsterpytestGitHub Actions

Kafka-PySpark-Delta Pipeline

Problem: Build scalable real-time streaming pipelines with exactly-once delivery guarantees.

Outcome: High-throughput data processing with exactly-once delivery and Delta checkpointing.

KafkaPySparkDelta LakePython

Cloud-Native Streaming Platform

Problem: Build multi-environment cloud infrastructure for real-time data processing.

Outcome: IaC-managed platform with exactly-once streaming and Delta Lake checkpointing.

AzureTerraformDatabricksDelta Lake

RAG Document QA Chatbot

Problem: Build intelligent document Q&A with retrieval-augmented generation.

Outcome: FastAPI + Streamlit app with vector search and configurable LLM backends.

FastAPIChromaDBRAGOpenAI
ETL Jobs

Enterprise Configuration-Driven Platform

Problem: Scale data pipeline development across enterprise business domains.

Outcome: Delivered numerous production jobs, improved deployment efficiency, maintained high availability, achieved comprehensive test coverage.

PySparkKafkaDelta LakeJSON Configs
Services

What I Can Help With

End-to-end data engineering capabilities from streaming pipelines to AI-ready platforms.

Data Engineering

ETL/ELT pipelines, data modeling, and data quality solutions.

PySparkSQLETL/ELTData Modeling

Kafka Streaming Pipelines

Real-time event streaming and exactly-once processing.

Apache KafkaStructured StreamingEvent Hubs

PySpark Engineering

Production-grade Spark applications with optimization.

PySparkApache SparkPerformance Tuning

Databricks Solutions

Lakehouse architecture and Delta Lake implementations.

DatabricksDelta LakeLakehouse

Terraform Infrastructure

Infrastructure as Code for data platforms.

TerraformAzureIaCDevOps

FastAPI Development

High-performance APIs and microservices.

FastAPIPythonREST APIsAsync

RAG Applications

Retrieval-augmented generation and vector search.

RAGVector DBsLLMsOpenAI

Data Platform Engineering

End-to-end data platform design and operations.

Data PlatformsCI/CDMonitoringTesting
Open Source Engineering

Production OSS Contributions

14+ Public Pull Requests across major open-source projects. I don't just build projects — I contribute to production software used by millions.

View GitHub Profile

Recent Contributions

urllib3 PR #5163

HTTP header bytes handling improvements.

View PR →

axios/axios#11113

Added missing fs import to README stream example.

View PR →

fastify/fastify#6880

Updated TypeScript docs to reference Fastify 5.x.

View PR →

strapi/strapi

Documentation and type improvements.

View Contributions →

trpc/trpc#7452

Secure error reporting documentation.

View PR →

vitest-dev/vitest

Testing framework improvements.

View Contributions →

jsdoc/jsdoc#2176

Align README Node.js requirement with package.json.

View PR →

Fastify

Web framework documentation updates.

View Contributions →
Skills

Tech Stack

Core technologies I use to build, test, and operate data platforms.

Data Engineering

PythonSQLPySparkApache SparkDelta LakeApache KafkaETL/ELTData ModelingSchema ValidationData Quality

Cloud & Data Platforms

DatabricksMicrosoft FabricAzure Event HubsADLS Gen2Delta Lake

AI / LLM

RAGLLMOpenAIChromaDBVector DatabasesSentence-TransformersFastAPIStreamlitGenAI

DevOps & Tools

GitGitHub ActionsGitLab CIDockerTerraformLinuxpytestJira
Certifications

Validated Technical Expertise

Industry-recognized credentials that back the skills and projects displayed here.

DP-700: Data Engineering with Microsoft Fabric

Microsoft · Dec 2025

Verify Credential →

Databricks PySpark Streaming Training

Accenture · Dec 2023 - Feb 2024

Google Data Analytics Professional

Coursera · 2023

Verify Credential →

NPTEL Management Information System (MIS)

Elite · 73% · IIT Kharagpur · Jul–Oct 2022

Verify Credential →
Blog

Engineering Notes

Deep dives into data platform engineering, streaming patterns, and career learnings.

Streaming Data Engineering

Building a Kafka → PySpark → Delta Pipeline

Hands-on guide to production PySpark Structured Streaming with Kafka, Delta Lake, and cloud deployment.

Read Article →
RAG & LLM

Building a RAG QA System with FastAPI and ChromaDB

How to build a retrieval-augmented generation chatbot with document ingestion, vector search, and LLM generation.

Read Article →
Certification

How I Passed the Databricks DE Associate Exam

My study plan, resources, tips, and what actually helped me pass the Databricks Data Engineer Associate exam.

Read Article →
Education

Academic Background

B.Tech in Computer Science and Technology

Madanapalle Institute of Technology & Science, Madanapalle, AP
2019–2023 · 7.97 CGPA · First Class Distinction

Intermediate

Sri Vivekananda Junior College, Chittoor, AP
2017–2019 · 9.67 CGPA

10th Grade

Z.P.H.S Kanipkam, Chittoor, AP
2016–2017 · 9.3 CGPA

JEE Mains

94.17 Percentile · 2019

Why Work With Me

Proven Production Impact

Quantifiable results and engineering discipline that delivers reliable data systems.

Python
Development
Automation
Engineering
Databricks
Certified Professional
14+
OSS Contributions
System
Integration
Production
Operations Experience
Contact

Currently Open To

✅ Python Engineer ✅ Automation Engineer ✅ Data Engineer ✅ System Integration Engineer ✅ Platform Engineering Roles

Location: Bengaluru, India

Remote: Open to Remote Opportunities

Reach out if you're hiring or want to collaborate on data platforms.

Email Me LinkedIn GitHub

Expected Response: Within 24 Hours