Hi There,
I'm Indir Lal
i am into
About Me
I am a results-driven Data Engineer specializing in designing and implementing high-performance, scalable ETL/ELT pipelines and cloud data architectures (Azure & GCP). With a proven track record of managing large-scale data systems—such as architecting a 60M+ record enterprise data warehouse in Google BigQuery from over 500M+ raw records—I excel at transforming complex, multi-source data streams into optimized, business-critical assets. Expert in Python, advanced SQL, PySpark, and modern data modeling practices, I build fault-tolerant data infrastructures that empower organizations with reliable, real-time insights and analytics-ready datasets.
Full-stack platform with PyTorch LSTM + XGBoost models and sentiment analysis. Improved Sharpe ratio 0.41 → 0.89 and cut max drawdown 34.2% → 19.4%.
Serverless ML pipeline forecasting air quality 72h ahead — feature store, ensemble models, SHAP explainability, scheduled GitHub Actions, and a live dashboard.
End-to-end pipeline via Azure Data Factory through ADLS and Databricks into a Medallion (Bronze/Silver/Gold) Lakehouse, with Terraform IaC, CI/CD, and Key Vault secrets.
LLM / retrieval (RAG) app answering questions over documents with citations, confidence scoring, and refusal behavior for low-confidence queries.
Jun 2026 - Present
Dec 2023 - Present