
Bhawandeep Singla
Delhi, India
Bhawandeep Singla
AI/LLM Integration & Backend Engineer — Python
Category : Artificial intelligence (AI)
I am a backend and distributed-systems engineer with 10 years of hands-on building, now specializing in adding production-grade AI/LLM features to existing products. I help teams ship the hard backend work — high-throughput services, low-latency APIs, reliable data systems — and wire in LLM capabilities (RAG, agents, developer tooling) that actually hold up in production rather than staying at the demo stage.
What I build for clients:
AI/LLM integration: RAG pipelines, multi-agent systems, and codebase-aware developer tooling on OpenAI and Anthropic models. I built a multi-agent, codebase-aware system spanning the full SDLC (exploration, design, coding, code review, testing) adopted across a 300-engineer org, and an AI on-call assistant ingesting alerts, deploys and runbooks that cut incident response time ~50%.
High-performance backends: Microservices and REST/event-driven APIs handling 4,000+ transactions/minute on platforms serving 25M+ users. I cut cart-API P95 latency from 1,900ms to 650ms and hotel-search query latency from 600ms to 40ms through query-plan analysis, index design and caching.
Reliability & observability: Raised service uptime from 99.01% to 99.99% with SLO tracking and error-budget practice; set up APM-driven alerting (Grafana, Sentry) that took mean-time-to-detect from hours to minutes.
Cloud & cost: AWS architecture (EC2, RDS, SQS, autoscaling) and Kubernetes; cut EC2 spend 40% via traffic-pattern-based autoscaling.
Data & resilience: Built data-replication and automatic primary→secondary failover across 120+ on-premises sites, lifting availability from ~95% to 99.95%.
From-scratch product builds: Built a B2B travel-booking portal (flights, hotels, activities) from the ground up, reaching ~8,000 daily users in year one.
Tech I work in:
Python, C++, Ruby on Rails · PostgreSQL, MongoDB, Redis, Elasticsearch · AWS, Kubernetes, GCP, CI/CD · microservices, REST & event/queue-driven design, system design · OpenAI & Anthropic APIs, agents, RAG, context engineering.
I work async-friendly, communicate proactively, and care about correctness, latency and cost — not just shipping something that runs once. Send me your problem and I will tell you honestly whether I'm the right fit and how I'd approach it.
What I build for clients:
AI/LLM integration: RAG pipelines, multi-agent systems, and codebase-aware developer tooling on OpenAI and Anthropic models. I built a multi-agent, codebase-aware system spanning the full SDLC (exploration, design, coding, code review, testing) adopted across a 300-engineer org, and an AI on-call assistant ingesting alerts, deploys and runbooks that cut incident response time ~50%.
High-performance backends: Microservices and REST/event-driven APIs handling 4,000+ transactions/minute on platforms serving 25M+ users. I cut cart-API P95 latency from 1,900ms to 650ms and hotel-search query latency from 600ms to 40ms through query-plan analysis, index design and caching.
Reliability & observability: Raised service uptime from 99.01% to 99.99% with SLO tracking and error-budget practice; set up APM-driven alerting (Grafana, Sentry) that took mean-time-to-detect from hours to minutes.
Cloud & cost: AWS architecture (EC2, RDS, SQS, autoscaling) and Kubernetes; cut EC2 spend 40% via traffic-pattern-based autoscaling.
Data & resilience: Built data-replication and automatic primary→secondary failover across 120+ on-premises sites, lifting availability from ~95% to 99.95%.
From-scratch product builds: Built a B2B travel-booking portal (flights, hotels, activities) from the ground up, reaching ~8,000 daily users in year one.
Tech I work in:
Python, C++, Ruby on Rails · PostgreSQL, MongoDB, Redis, Elasticsearch · AWS, Kubernetes, GCP, CI/CD · microservices, REST & event/queue-driven design, system design · OpenAI & Anthropic APIs, agents, RAG, context engineering.
I work async-friendly, communicate proactively, and care about correctness, latency and cost — not just shipping something that runs once. Send me your problem and I will tell you honestly whether I'm the right fit and how I'd approach it.
Portfolio
Working hours
- Monday:08h00 To 18h00
- Tuesday:08h00 To 18h00
- Wednesday:08h00 To 18h00
- Thursday:08h00 To 18h00
- Friday:08h00 To 18h00
- Saturday:Not available
- Sunday:Not available
Please sign in as a customer to give your feedback





