Presentation
Portfolio
Background
Impact: Supported training, validation, and benchmarking of advanced AI systems through high-quality, structured task execution.
Focus Areas: Machine Learning, Statistics, Data Analysis, Python, LLM Evaluation
Sustainability Intelligence Pipeline: Built an NLP-driven scoring and classification pipeline to extract structured sustainability signals from unstructured public datasets and documents.
Impact: Enabled automated generation of structured sustainability insights for downstream decision workflows.
Techniques: NLP pipelines, semantic filtering, rule-based + model-based scoring
Stack: Python, NLP, Data Processing
End-to-End ML System: Designed and implemented a hybrid ML + rule-based + LLM system for structured scoring from noisy textual data.
Impact: Delivered production-ready outputs aligned with business decision workflows.
Techniques: Feature engineering, hybrid modeling, structured output design
ML Pipeline & MLOps Integration: Implemented ML pipeline improvements and experiment tracking for model evaluation and reproducibility.
Impact: Improved model validation reliability and streamlined experimentation workflows.
Techniques: Model evaluation, pipeline structuring, experiment tracking
Architecting collaborative ML platform for multi-trillion parameter models (WIP)
Coordinating with MLOps and ML Architecture experts
Impact: Improved scoring accuracy by ~15%, reduced manual QA workload, and scaled automated auditing across high call volumes.
Techniques: Feature engineering, supervised modeling, hyperparameter tuning
Stack: Python, scikit-learn, Librosa, HuggingFace, TensorFlow/PyTorch, AWS S3, Pandas, NumPy
Speech Emotion Recognition (SER): Developed an SER workflow using MFCC and spectrogram-based feature extraction to classify tonal emotion from call recordings.
Impact: Enabled emotion-aware scoring, improving depth and reliability of call quality assessments.
Techniques: Signal processing, classification, feature extraction
Stack: PyTorch, TensorFlow, Librosa, NumPy, Pandas
Multilingual NLP Pipelines: Built end-to-end pipelines for sentiment analysis, emotion detection, topic modeling, and summarization on multilingual call transcripts.
Impact: Transformed large volumes of unstructured conversation data into structured, actionable insights for product and CX teams.
Techniques: LDA, BERTopic, NMF, transformer-based summarization (T5, BART, GPT-2)
Stack: HuggingFace, SpaCy, Gensim, Sumy, NLTK
Automated Model Regression & QA Validation: Built batch inference pipelines and regression testing frameworks to validate ML model performance across releases.
Impact: Reduced manual QA cycles and improved reliability and consistency of ML deployments.
Techniques: Batch evaluation, metric tracking
Stack: Python, PyTest, Pandas
Impact: Reduced manual review time and improved accessibility of insights for non-technical stakeholders.
Techniques: Seq2Seq, TextRank, transformer summarization
Stack: HuggingFace Transformers, T5/BART, Python, SpaCy, Gensim, NLTK
Built MVP analytics reporting layer for a new product offering to validate use cases and demonstrate value to early customers.
Impact: Delivered working MVP within 2 months, contributing to initial customer acquisition.
Techniques: EDA, metric design, aggregation, rapid prototyping
Stack: Python, PostgreSQL, Pandas, NumPy
Developed customer journey analytics and automated reporting pipelines for Business Development and Customer Success teams.
Impact: Improved visibility into funnel performance and supported better retention strategies.
Techniques: Funnel analysis, cohort-style analysis, KPI design
Stack: Python, PostgreSQL, Pandas, NumPy
Designed and implemented predictive lead scoring models (US real estate) using behavioral and geo-spatial features.
Impact: Improved targeting efficiency and downstream conversion likelihood.
Techniques: Classification, feature importance analysis, geospatial modeling
Stack: XGBoost, scikit-learn, SQL, Pandas, NumPy
Built ticket clustering and root-cause analysis system to identify recurring issues and patterns in support data.
Impact: Enabled data-driven prioritization of product fixes and operational improvements.
Techniques: KMeans, DBSCAN, PCA, EDA, feature engineering
Stack: scikit-learn, Pandas, NumPy, Python
Impact: Improved decision-making for acquisition and retention strategies.
Stack: SQL, Python, Google Sheets, Data Studio, Metabase
Built lead qualification scoring system for US real estate workflows.
Impact: Streamlined lead prioritization and improved sales efficiency.
Techniques: Rule-based scoring, exploratory data analysis, feature identification
Stack: Python, SQL, Pandas
Impact: Improved early-stage lead filtering and laid the foundation for later predictive lead scoring systems.
Techniques: Rule-based scoring, exploratory data analysis, feature identification
Stack: Python, SQL, Pandas
Performed exploratory data analysis and statistical validation to identify patterns in customer behavior and conversion trends.
Impact: Provided early insights that informed feature design for downstream ML models.
Techniques: Hypothesis testing, statistical analysis, data exploration

