San Francisco, CA · Open to DevOps & platform roles
Anvar Salvar
Senior DevOps Engineer
Six years building AWS infrastructure, Kubernetes platforms and CI/CD
automation. I build the shared delivery tooling that 18 production
services across 7 application teams ship through — EKS,
Terraform, Helm, GitHub Actions and Argo CD.
And I build with AI rather than just using it: Claude Code Agent Skills
and MCP integrations for infrastructure, deployment validation and
troubleshooting.
One deployment path for 18 services, instead of 18 copies of one
Shared delivery tooling for 18 production services across 7 application
teams. Reworked CI/CD to build the artifact once and promote that same
image through every environment, with consistent approval and rollback
behaviour across all of them.
~40% faster
deployment lead time
GitHub Actions
Argo CD
Helm
ECR
OIDC
Infrastructure
DNS, TLS and secrets without opening a ticket
Replaced manual provisioning with declarative resources the application
team owns in its own repo. Standard service setup dropped from about
three days across three queues to a single pull request.
~3 days → <4 hours
standard service setup
cert-manager
ExternalDNS
Secrets Store CSI
Gateway API
Helm
Verification
A deploy isn’t finished because the pipeline went green
Added deployment safeguards that catch AWS authentication failures,
incomplete rollouts and incorrect releases before they reach
production — rather than after someone notices in the dashboard.
The principle underneath all of it: the only trustworthy signal is the
one that ties the running workload back to the artifact you actually
built. A pipeline that reports success without checking that is
measuring the service, not the release.
rollout verification
digest checks
OIDC preflight
rollback automation
Reliability
Fewer pages, without losing the ones that matter
Built the Prometheus/Grafana observability baseline with SLO-based alerts
and Alertmanager routing. Earlier, at AppsFlyer, I tuned alerting for
Kubernetes and Kafka workloads down from ~30–40 pages a week to
about 15 — without losing detection of real incidents, which is
the only version of that number worth quoting.
~35% fewer
non-actionable production alerts
Prometheus
Grafana
Alertmanager
SLOs
production on-call
Scale
Terraform modules, autoscaling, and $18K a month back
Reusable Terraform for AWS and EKS environments used by three application
teams took new-environment provisioning from ~2 days to under 2 hours.
Rightsizing and autoscaling across six processing services cut EKS
worker-node spend while holding latency and throughput targets.
~$18K/mo (~15%)
EKS worker-node spend removed
Terraform
EKS
autoscaling
Kafka
capacity
Intelligence
Agent Skills and MCP, with the boundary enforced in Kubernetes
Claude Code Agent Skills give repeatable Terraform, deployment-validation
and troubleshooting workflows a fixed shape instead of an open prompt, and
MCP integrations connect the nine-person platform team to GitHub,
Kubernetes, AWS and operational tooling.
Reads run locally per engineer; anything that changes infrastructure goes
through one path, with a person approving the change. The boundary is
RBAC and quotas, not wording.
Claude Code
Agent Skills
MCP
RBAC
human-in-the-loop
Stack
Cloud & Kubernetes
AWSEKSEC2IAMVPCRoute 53RDSECRDockerLinux
Infrastructure & platform
TerraformHelmGateway APIcert-managerExternalDNSSecrets Store CSI
I’m looking for a DevOps or platform team that takes fundamentals
seriously and is genuinely putting AI into how engineers work — with
the guardrails to do it safely. Based in San Francisco, open to relocation.