
Abhinaba Banerjee
Vijayawada, India
Abhinaba Banerjee
Generative AI & RAG Specialist
Category : Artificial intelligence (AI)
I am an AI/ML Engineer specializing in Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) systems, and inference performance optimization.
I help startups, developers, and businesses build production-ready AI applications, fix broken LLM code, and optimize local or cloud model deployments for speed and cost efficiency.
How I Can Help You:
• Custom RAG Development: Connect LLMs to your private business data, PDFs, or databases with high precision and low hallucination.
• Inference & Cost Optimization: Accelerate local/cloud LLM throughput, reduce API costs, and tune latency using modern frameworks (vLLM, Ollama, CUDA optimization).
• Code Debugging & Unblocking: Fix bugs and logical bottlenecks in your LangChain, LlamaIndex, or custom Python ML scripts.
• AI Agents & Integrations: Implement function calling, structured outputs, and automated workflows into your existing software.
Tech Stack & Tools:
Python, PyTorch, LangChain, LlamaIndex, Hugging Face, vLLM, CUDA, Vector Databases (Pinecone, Qdrant, Chroma, FAISS).
Whether you need a quick 1-hour code debugging session, a prototype built, or an end-to-end RAG pipeline implemented, feel free to reach out to discuss your project!
I help startups, developers, and businesses build production-ready AI applications, fix broken LLM code, and optimize local or cloud model deployments for speed and cost efficiency.
How I Can Help You:
• Custom RAG Development: Connect LLMs to your private business data, PDFs, or databases with high precision and low hallucination.
• Inference & Cost Optimization: Accelerate local/cloud LLM throughput, reduce API costs, and tune latency using modern frameworks (vLLM, Ollama, CUDA optimization).
• Code Debugging & Unblocking: Fix bugs and logical bottlenecks in your LangChain, LlamaIndex, or custom Python ML scripts.
• AI Agents & Integrations: Implement function calling, structured outputs, and automated workflows into your existing software.
Tech Stack & Tools:
Python, PyTorch, LangChain, LlamaIndex, Hugging Face, vLLM, CUDA, Vector Databases (Pinecone, Qdrant, Chroma, FAISS).
Whether you need a quick 1-hour code debugging session, a prototype built, or an end-to-end RAG pipeline implemented, feel free to reach out to discuss your project!
Working hours
- Monday:08h00 To 18h00
- Tuesday:08h00 To 18h00
- Wednesday:08h00 To 18h00
- Thursday:08h00 To 18h00
- Friday:08h00 To 18h00
- Saturday:Not available
- Sunday:Not available
- 🇬🇧 English
- 🇫🇷 French
- 🇮🇳 Hindi
Please sign in as a customer to give your feedback


