Presentation
What I do:
- Design and hands-on experience in creating and operating all the infrastructure needed for production, and help the software engineers team design and implement scalable/reliable solutions.
- Ship reliable, observable services (Prometheus, Grafana, Loki, OpenTelemetry) with SLOs & actionable alerts.
- Balance security, performance, and cloud cost; drive FinOps and incident response playbooks.
- Lead cross-functional teams, hiring/coaching engineers, and improving delivery flow.
Highlights
- Real-time data platform serving 1,000+ machines and 10k+ daily users, with resilient ingestion and processing pipelines.
- Built CI/CD from scratch for a monorepo (10 services): artifact build <10 min, prod deploy <5 min, near-zero downtime releases.
- Delivered full-stack observability using the Grafana LGTM stack (Loki, Grafana, Tempo, Mimir) with SLOs, runbooks, and on-call readiness.
Outside work, I enjoy learn new technologies, which help me drive technical innovation through open source contributions, grounded in deep study, and amplified by teaching.

