Senior Site Reliability Engineer
Juul Labs
4h ago
0$185k - $227kDevopsUnited Stateshimalayas
Site-Reliability-EngineeringInfrastructure-EngineeringCloud-EngineerSystems-EngineeringDevOps-EngineerSenior
Job Description
THE COMPANY:Juul Labs's mission is to transition the world’s billion adult smokers away from combustible cigarettes, eliminate their use, and combat underage usage of our products. We have the opportunity to address one of the world’s most intractable challenges through a commitment to exceptional quality, research, design, and innovation. Backed by leading technology investors, we are committed to the same excellence when it comes to hiring great talent.We are a diverse team that is united by this common purpose and we are hiring the world’s best engineers, scientists, designers, product managers, operations experts, and customer service and business professionals. If the opportunity to build your career is compelling, read on for more details.ROLE AND RESPONSIBILITIES:A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance of Juul’s hybrid cloud infrastructure (Nutanix, AWS/GCP). This involves leading automation efforts, architecting for reliability, and acting as the final escalation point for critical incidents to ensure the platform is scalable and efficient.Nutanix Platform ManagementDesign, deploy, and maintain enterprise-scale Nutanix AHV clusters and Prism Central for multi-cluster managementExpert-level proficiency with Nutanix CLI (nCLI and acli) for advanced operations, troubleshooting, and automationDevelop automation scripts using Nutanix REST APIs, Python SDK, PowerShell, and Terraform for infrastructure-as-codeCreate and manage VM templates, golden images, and standardized deployment catalogs for consistent provisioningDesign disaster recovery solutions using Leap, Protection Domains, cross-cluster replication, and metro clusteringImplement network micro-segmentation using Nutanix Flow and configure RBAC, encryption, and security hardeningLead L3 troubleshooting using advanced diagnostics, log analysis (CVM, Genesis), NCC health checks, and cluster service resolutionConfigure high availability, VM affinity rules, QoS policies, and optimize performance for mission-critical workloadsManage AHV networking with OVS bridges, VLANs, bonds, LACP and implement resource reservations and workload balance.Design, deploy, and maintain hybrid cloud infrastructure across Nutanix HCI, AWS, and GCP platformsArchitect and implement multi-cloud solutions ensuring high availability, scalability, and disaster recoveryCloud Platform EngineeringArchitect and deploy enterprise-scale, highly available multi-cloud solutions across AWS and GCP with multi-region/multi-account strategiesExpert-level proficiency with AWS CLI, GCP CLI, SDK, boto3, and Python for advanced automation and infrastructure orchestrationDesign AWS Organizations and GCP Organization hierarchies with consolidated billing, IAM policies, and centralized governanceConfigure and manage AWS Systems Manager (SSM) including Session Manager, Run Command, State Manager, and Automation for centralized fleet operationsImplement centralized logging using CloudWatch/CloudTrail and GCP Cloud Logging with S3/Cloud Storage aggregationIntegrate AWS and GCP with Splunk using HEC, CloudWatch subscriptions, Pub/Sub, Dataflow, and cloud-specific add-ons for SIEM correlationDesign and deploy advanced load balancing solutions with AWS ALB/NLB/ELB and GCP Cloud Load Balancing including SSL termination and auto-scalingDevelop infrastructure-as-code using Terraform, CloudFormation, CDK for repeatable multi-cloud deployments and CI/CD pipelinesConfigure AWS SSO, cross-account IAM roles, GCP Workload Identity, and federated access for centralized identity managementDesign VPC architectures with AWS Transit Gateway/PrivateLink and GCP Shared VPC/VPC peering for hybrid connectivityManage containerized workloads using EKS, GKE, ECS, Cloud Run with service mesh, observability, and security best practicesImplement disaster recovery using AWS Backup, Cross-Region Replication, GCP snapshots, and multi-region failover strategiesLead L3 troubleshooting using CloudWatch Insights, GCP Cloud Trace, VPC Flow Logs, X-Ray, and vendor support escalationPerform cost optimization through Reserved Instances, Committed Use Discounts, rightsizing, and automated resource lifecycle managementSystem AdministrationAdminister and support Windows Server and Unix/Linux environments in production and non-production settingsPerform OS-level hardening, patch management, and security compliance across heterogeneous systemsAutomate routine administrative tasks using PowerShell, Bash, Python, or similar scripting languagesManage GitHub organization settings, user permissions, repository access controls, and monitor GitHub Actions workflows and repository health across multiple teamsConfigure Splunk forwarders, heavy forwarders and other integrations for data ingestion from cloud and on-premises sourcesPERSONAL AND PROFESSIONAL QUALIFICATIONS:8-12+ years infrastructure experience with 8+ years in Nutanix HCI and enterprise cloud AWS/GCP)Expert-level skills in Pyt
