Ratnapriya Lal

Ratnapriya Lal

Ahmedabad, India

Ratnapriya Lal

Senior Data Scientist - ML, NLP, AI
Data Scientist with 6+ years of experience building practical ML and NLP systems for real-world use cases. I’ve worked across customer analytics, conversational data, and decision-support workflows, focusing on turning messy, unstructured data into reliable, structured outputs. I work primarily with Python-based pipelines combining statistical models, feature engineering, and LLM-driven approaches where needed. Comfortable handling ambiguity, noisy data, and evolving product requirements, with a focus on building solutions that are actually usable in production.

Working hours

  • Monday:10h00 To 20h00
  • Tuesday:10h00 To 20h00
  • Wednesday:10h00 To 20h00
  • Thursday:10h00 To 20h00
  • Friday:10h00 To 20h00
  • Saturday:10h00 To 20h00
  • Sunday:10h00 To 20h00
Applied ML & LLM Systems (SuperAnnotate, Alignerr): Contributed to large-scale ML and LLM evaluation pipelines, solving data science, statistical reasoning, and Python-based tasks across diverse problem domains.
Impact: Supported training, validation, and benchmarking of advanced AI systems through high-quality, structured task execution.
Focus Areas: Machine Learning, Statistics, Data Analysis, Python, LLM Evaluation

Sustainability Intelligence Pipeline: Built an NLP-driven scoring and classification pipeline to extract structured sustainability signals from unstructured public datasets and documents.
Impact: Enabled automated generation of structured sustainability insights for downstream decision workflows.
Techniques: NLP pipelines, semantic filtering, rule-based + model-based scoring
Stack: Python, NLP, Data Processing

End-to-End ML System: Designed and implemented a hybrid ML + rule-based + LLM system for structured scoring from noisy textual data.
Impact: Delivered production-ready outputs aligned with business decision workflows.
Techniques: Feature engineering, hybrid modeling, structured output design

ML Pipeline & MLOps Integration: Implemented ML pipeline improvements and experiment tracking for model evaluation and reproducibility.
Impact: Improved model validation reliability and streamlined experimentation workflows.
Techniques: Model evaluation, pipeline structuring, experiment tracking
Leading team for semi-assisted context classification of Vedic texts
Architecting collaborative ML platform for multi-trillion parameter models (WIP)
Coordinating with MLOps and ML Architecture experts
Dynamic Call Quality Audit System: Engineered and deployed an ML-driven call-quality scoring system using audio + transcript pipelines, integrating Speech Emotion Recognition (SER) and NLP models for multi-parameter evaluation.
Impact: Improved scoring accuracy by ~15%, reduced manual QA workload, and scaled automated auditing across high call volumes.
Techniques: Feature engineering, supervised modeling, hyperparameter tuning
Stack: Python, scikit-learn, Librosa, HuggingFace, TensorFlow/PyTorch, AWS S3, Pandas, NumPy

Speech Emotion Recognition (SER): Developed an SER workflow using MFCC and spectrogram-based feature extraction to classify tonal emotion from call recordings.
Impact: Enabled emotion-aware scoring, improving depth and reliability of call quality assessments.
Techniques: Signal processing, classification, feature extraction
Stack: PyTorch, TensorFlow, Librosa, NumPy, Pandas

Multilingual NLP Pipelines: Built end-to-end pipelines for sentiment analysis, emotion detection, topic modeling, and summarization on multilingual call transcripts.
Impact: Transformed large volumes of unstructured conversation data into structured, actionable insights for product and CX teams.
Techniques: LDA, BERTopic, NMF, transformer-based summarization (T5, BART, GPT-2)
Stack: HuggingFace, SpaCy, Gensim, Sumy, NLTK

Automated Model Regression & QA Validation: Built batch inference pipelines and regression testing frameworks to validate ML model performance across releases.
Impact: Reduced manual QA cycles and improved reliability and consistency of ML deployments.
Techniques: Batch evaluation, metric tracking
Stack: Python, PyTest, Pandas
Transcript Summarization for CX Insights: Designed extractive and abstractive summarization workflows to condense long conversations into high-signal summaries.
Impact: Reduced manual review time and improved accessibility of insights for non-technical stakeholders.
Techniques: Seq2Seq, TextRank, transformer summarization
Stack: HuggingFace Transformers, T5/BART, Python, SpaCy, Gensim, NLTK

Built MVP analytics reporting layer for a new product offering to validate use cases and demonstrate value to early customers.
Impact: Delivered working MVP within 2 months, contributing to initial customer acquisition.
Techniques: EDA, metric design, aggregation, rapid prototyping
Stack: Python, PostgreSQL, Pandas, NumPy

Developed customer journey analytics and automated reporting pipelines for Business Development and Customer Success teams.
Impact: Improved visibility into funnel performance and supported better retention strategies.
Techniques: Funnel analysis, cohort-style analysis, KPI design
Stack: Python, PostgreSQL, Pandas, NumPy

Designed and implemented predictive lead scoring models (US real estate) using behavioral and geo-spatial features.
Impact: Improved targeting efficiency and downstream conversion likelihood.
Techniques: Classification, feature importance analysis, geospatial modeling
Stack: XGBoost, scikit-learn, SQL, Pandas, NumPy

Built ticket clustering and root-cause analysis system to identify recurring issues and patterns in support data.
Impact: Enabled data-driven prioritization of product fixes and operational improvements.
Techniques: KMeans, DBSCAN, PCA, EDA, feature engineering
Stack: scikit-learn, Pandas, NumPy, Python
Developed CRM analytics and automated reporting systems to track customer lifecycle and engagement.
Impact: Improved decision-making for acquisition and retention strategies.
Stack: SQL, Python, Google Sheets, Data Studio, Metabase

Built lead qualification scoring system for US real estate workflows.
Impact: Streamlined lead prioritization and improved sales efficiency.
Techniques: Rule-based scoring, exploratory data analysis, feature identification
Stack: Python, SQL, Pandas
Built an initial lead qualification scoring system to prioritize high-intent leads in US real estate workflows.
Impact: Improved early-stage lead filtering and laid the foundation for later predictive lead scoring systems.
Techniques: Rule-based scoring, exploratory data analysis, feature identification
Stack: Python, SQL, Pandas

Performed exploratory data analysis and statistical validation to identify patterns in customer behavior and conversion trends.
Impact: Provided early insights that informed feature design for downstream ML models.
Techniques: Hypothesis testing, statistical analysis, data exploration
  • 🇬🇧 English
  • 🇮🇳 Hindi
0.0 (0)
0
0
0
0
0
⚠️
Please sign in as a customer to give your feedback

    Need a service ?

    Do you want to hire the services of this Professional ?

    Post your need and you will receive dozens of proposals from the Professionals of the community.