Harry Evans
Dublin, Ireland
Harry Evans
Software Engineer
Category : Web development
I'm a Senior Full-Stack AI Engineer with over 8 years of experience, specializing in large-scale AI/LLM serving infrastructure for the past 2.5+ years. My expertise spans model serving and inference optimization, streaming APIs, multi-region deployment, and client-side integration across React and TypeScript. At ElevenLabs, I owned the path from model abstraction through to playback, cutting perceived latency 20-40% for international users and helping the platform scale to tens of millions of daily clips without queue collapse. At Microsoft, I led reliability engineering on the Azure OpenAI serving path, eliminating a multi-region retry-storm pattern and improving p99 latency 45-65% under load. I work across the full stack, from GPU capacity management and continuous batching to the frontend code that makes streaming responses feel instant, partnering with research, infrastructure, and product teams to turn competing latency and scalability demands into serving patterns that actually ship.
Working hours
- Monday:08h00 To 18h00
- Tuesday:08h00 To 18h00
- Wednesday:08h00 To 18h00
- Thursday:08h00 To 18h00
- Friday:08h00 To 18h00
- Saturday:Not available
- Sunday:Not available
Please sign in as a customer to give your feedback


