Designing and operating distributed storage clusters using Ceph and S3-compatible object storage. Experienced in data replication, fault tolerance, and maintaining high availability across multi-node environments under production workloads.
Applying SRE principles to reduce toil and improve system reliability — including defining SLOs, building alerting pipelines, performing root cause analysis, and implementing runbooks that cut incident response time.
Deploying and managing containerized workloads on Kubernetes with a focus on scalability and resilience. Experienced with Helm, resource optimization, persistent storage integration, and cluster observability using Prometheus and Grafana.