Job Title: Senior AI-Enabled Platform / SRE Engineer.
Location: Dallas, TX or Scottsdale, AZ (Hybrid).
Employment Type: Contract (W2 through ZipStaff).
About the Opportunity:
ZipStaff is seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills to support a major healthcare organization.
This hybrid role (Dallas, TX or Scottsdale, AZ) focuses on platform reliability at scale and using AI/LLMs to automate operations. The initial assignment is approximately one year, with potential extension.
What You’ll Do:
- Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
- Leverage Generative AI (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
- Implement API and microservices reliability solutions using Apigee / Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
- Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
- Support active-active deployments, disaster-recovery readiness, and multi-datacenter Kubernetes environments.
- Develop observability and monitoring using Splunk, Grafana, Datadog, and AppDynamics.
- Drive SRE best practices by partnering with cross-functional teams on reliability, security, incident management, and continuous improvement.
Required Qualifications:
- 5+ years of strong hands-on Kubernetes platform experience with GKE and Rancher RKE2, including multi-cluster management, troubleshooting, and performance optimization.
- 5+ years of advanced programming in Python and Java (Node.js preferred for integrations and automation).
- Strong SRE background: reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
- Strong experience in GCP, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
- Hands-on observability experience with Splunk, Grafana, Datadog, and/or AppDynamics.
- Experience with API and microservices reliability (Apigee / Apigee X, REST, GraphQL, traffic routing, canary, failover).
- Experience applying LLMs (Gemini, Llama, Mistral, Qwen, or similar) to alert analysis, incident triage, automation, or operational workflows.
- Ability to work hybrid in Dallas, TX or Scottsdale, AZ.
- Must be legally authorized to work in the United States without sponsorship now or in the future.
Preferred Qualifications:
- Experience supporting highly available, multi-datacenter production platforms.
- Prior healthcare or regulated-enterprise SRE experience.
- Demonstrated AIOps / GenAI-for-operations implementations in production.
About ZipStaff:
ZipStaff partners with leading organizations to connect skilled technology professionals with high-impact contract opportunities. We focus on quality matches and long-term success.
To apply, please submit your resume highlighting your GKE / RKE2, Python/Java automation, observability, and any AI-driven operations work.