Job Title: Site Reliability Engineer (SRE).
Location: Scottsdale / Shea, AZ (Hybrid).
Employment Type: Contract (W2 through ZipStaff).
About the Opportunity:
ZipStaff is seeking a Site Reliability Engineer to support large-scale enterprise applications for a major healthcare organization.
This is a hybrid role based at the Shea, AZ location and focused on cloud-native operations, observability, automation, and production support across hybrid (on-premises and cloud) environments. Immediate availability is preferred.
What You’ll Do:
- Support production operations and reliability for large-scale, high-performance applications.
- Develop automation scripts and build APM dashboards to monitor end-to-end transaction journeys.
- Implement observability using OpenTelemetry (OTEL), distributed tracing, monitoring, and incident management.
- Manage containerized applications in Kubernetes environments (GKE, RKE, or AKS).
- Support cloud migration and containerization initiatives (GCP, AWS, Azure, Rancher, OpenShift, or similar).
- Participate in 24x7 on-call rotations and meet incident-response SLAs.
- Troubleshoot distributed systems, networking, and API gateway issues.
- Partner with engineering and operations teams to improve reliability, automation, and operational excellence.
Required Qualifications:
- 5 years of experience in Site Reliability Engineering, Production Operations, or Platform Engineering supporting large-scale applications in hybrid environments.
- 5 years developing automation scripts and building APM dashboards for end-to-end transaction monitoring.
- 3+ years of hands-on programming in one or more of: Go, Python, Java, or Rust.
- Working knowledge of relational and/or NoSQL databases (Oracle, SQL Server, PostgreSQL, MongoDB, Redis, ClickHouse, PL/SQL, or time-series databases).
- Experience with cloud migration and containerization (GCP, AWS, Azure, Rancher, OpenShift, or similar).
- Experience managing containers in Kubernetes (GKE, RKE, or AKS).
- Strong experience implementing observability (OTEL, distributed tracing, monitoring, incident management).
- Strong networking fundamentals (TCP/IP, HTTP, DNS, load balancing, service mesh).
- Experience participating in 24x7 on-call rotations and meeting incident SLAs.
- Ability to work hybrid in Scottsdale / Shea, AZ.
- Must be legally authorized to work in the United States without sponsorship now or in the future.
Preferred Qualifications:
- Experience managing highly available, customer-facing platforms.
- Hands-on experience with Splunk, Dynatrace, AppDynamics, Grafana, and/or Prometheus.
- Experience with CI/CD and Agile tools (Rally, Confluence, DevOps platforms).
- Knowledge of in-memory caching, especially Redis.
- Strong troubleshooting across distributed systems and API gateways.
- Experience with Google Cloud services (GCS, Cloud SQL, Spanner, BigQuery).
- Experience supporting HashiCorp Vault.
- Exposure to Vertex AI, Generative AI, or cloud analytics platforms.
- Familiarity with GraphQL frameworks (Apollo, Prisma, or Hasura).
About ZipStaff:
ZipStaff partners with leading organizations to connect skilled technology professionals with high-impact contract opportunities. We focus on quality matches and long-term success.
To apply, please submit your resume highlighting your SRE / production operations experience, Kubernetes, observability, and on-call support background.