Utkarsh Babbar
I keep systems reliable at scale.
Member of Technical Staff (SDE2) on Platform & DevOps at Salesforce. I build high-throughput telemetry systems, run Kubernetes at scale, and turn operational chaos into automated, observable infrastructure.
About
I'm a Platform & DevOps engineer at Salesforce, where I build and operate the telemetry platforms that keep large-scale production systems reliable. My work spans Splunk, Grafana, Prometheus, and OpenTelemetry pipelines, SLA/SLO/SLI engineering, and high-severity incident response.
Day-to-day I write high-throughput telemetry services in Go and Python, orchestrate clusters on AWS EKS with Docker and Helm, and provision infrastructure as code with Terraform and Ansible. I care about the details that make infrastructure boring in the best way: predictable deploys, clean pipelines, and telemetry you can trust when things break.
Before focusing on platform and observability, I worked across platform and infrastructure systems and cloud automation — which still shapes how I build tools that solve real problems for the engineers who use them.
Experience
Member of Technical Staff (SDE2) — Platform & DevOps
- Designed asynchronous, high-throughput cloud telemetry services in Go and Python, optimizing back-end pipelines to scale reliably across production environments.
- Primary system owner for scalable Splunk, Grafana, and Prometheus log pipelines, ingestion routing, and alerting-rule topologies across enterprise clusters.
- Engineered real-time telemetry dashboards against strict SLAs/SLOs/SLIs, cutting production mean-time-to-detection (MTTD) by 30% for core infrastructure failures.
- Orchestrated high-volume clustered nodes on AWS EKS with Docker — clean Helm charts and custom ingress components for massive transaction footprints.
- Led incident-response rotations for high-severity (Sev 1/2) issues, establishing RCA records and automated runbooks to reduce MTTR by 20%.
Software Development Engineer 1 — Platform & Infrastructure
- Built decentralized telemetry environments — structured logging and metrics via OpenTelemetry — and identified scale boundaries inside core clusters.
- Refactored high-load microservices in C++ and Java, reducing platform memory consumption by 15% while sustaining a 20% increase in active workloads.
- Provisioned scalable CI architectures on AWS (Lambda, API Gateway, DynamoDB) from scratch via modular, reusable Terraform and Ansible.
- Developed Python and Go automation tooling wrapped in Molecule test layers, cutting execution latency by 40%.
- Configured end-to-end monitoring with Prometheus and Grafana across 30+ services; mentored 4 junior engineers on production cloud deployments.
Software Development Intern — Cloud Automation
- Crafted multi-cloud assets with Terraform, Terragrunt, and Python, reducing infrastructure rollout overhead by 40%.
- Deployed application collection and metric-visualization architectures with Prometheus and Grafana across 50+ microservice pods, cutting debugging time by 25%.
Stack
Observability
Containers & Orchestration
Infrastructure as Code
Cloud & Architecture
Reliability & Ops
Languages
Education
B.E. — Electronics & Computer Engineering
Indian School Certificate (ISC), Class XII
Awards & Certifications
SAFe Agilist (Scrum) — earned with a 90% score.
Amadeus Game Changer Award — high-impact cloud architecture during complex migrations.
AWS Certified Cloud Practitioner (CCP) & Microsoft Azure Fundamentals (AZ-900).
2nd of 150+ teams at HackOWASP; 650+ problems solved on InterviewBit & LeetCode.
Contact
Let's build reliable systems.
Open to conversations about platform engineering, observability, and DevOps. The fastest way to reach me is email.