Building Real-Time Streaming & AI-Ready Data Systems with PySpark, Kafka, Delta Lake, and Databricks.
I specialize in turning high-velocity event streams into reliable, AI-ready data products. With hands-on experience across the full data lifecycle, I design, build, and support production-grade data platforms.
I am a Data & AI Platform Engineer with 2+ years validating, building, and operating production-grade data platforms and lakehouse systems at Accenture. Contributed to 50+ data products and delivered 90+ production PySpark ETL jobs while maintaining 99.5% pipeline uptime and 95%+ test coverage. I focus on configuration-driven pipelines, data quality, testing, and reliable production operations.
My core focus is on configuration-driven data platforms, real-time pipelines, lakehouse architectures, and the engineering discipline that makes production systems reliable — testing, CI/CD, observability, and security.
# Platform in a nutshell data_source = "Kafka" processing = "PySpark Structured Streaming" storage = "Delta Lake" ci_cd = "GitLab CI/CD" testing = "pytest, 95%+ coverage" platform = "AI-ready, scalable, secure"
Hands-on development and validation of configuration-driven data pipelines serving end-to-end supply-chain data products.
A blend of enterprise-scale work and personal engineering labs, each focused on real-world reliability, performance, and impact.
PySpark, Kafka, JSON-based job definitions, DDL, and SQL validation for Accenture data products.
Azure Event Hubs, Databricks, ADLS Gen2, and Terraform. Multi-environment IaC and exactly-once streaming with Delta Lake checkpointing.
FastAPI + Streamlit app with ChromaDB, sentence-transformers, and OpenAI/LLaMA backends. Response caching and configurable LLM backends.
Production-style streaming ingestion, schema enforcement, watermark-based deduplication, and exactly-once checkpointing. Benchmarked at 31k–45k rows/sec on a 4-core laptop.
Verified upstream contributions, active PRs, and maintainer-style reviews across well-known open-source projects. Last verified: 2026-07-31.
$35 bounty: coordinate auth token refresh across tabs.
View PR →Maintainer-style review covering code, tests, docs, and conventions.
View Review →Core technologies I use to build, test, and operate data platforms.
Industry-recognized credentials that back the skills and projects displayed here.
Accenture · Dec 2023 - Feb 2024
Elite · 73% · IIT Kharagpur · Jul–Oct 2022
Verify Credential →Deep dives into data platform engineering, streaming patterns, and career learnings.
Hands-on guide to production PySpark Structured Streaming with Kafka, Delta Lake, and cloud deployment.
Read Article →How to build a retrieval-augmented generation chatbot with document ingestion, vector search, and LLM generation.
Read Article →My study plan, resources, tips, and what actually helped me pass the Databricks Data Engineer Associate exam.
Read Article →Madanapalle Institute of Technology & Science, Madanapalle, AP
2019–2023 · 7.97 CGPA · First Class Distinction
Sri Vivekananda Junior College, Chittoor, AP
2017–2019 · 9.67 CGPA
Z.P.H.S Kanipkam, Chittoor, AP
2016–2017 · 9.3 CGPA
94.17 Percentile · 2019
Open to Data & AI Platform Engineer, Data Engineer, and Cloud Data Engineering roles in Bengaluru and remote. Reach out if you are hiring or want to collaborate.
sasidharmopuru@gmail.comI typically respond within 24–48 hours.