Senior Site Reliability Engineer
SunCore Digital
4h ago
0$145k - $185kDevopsUnited Kingdom, United Stateshimalayas
Site-Reliability-EngineeringDevOpsInfrastructure-EngineeringProduction-EngineeringPlatform-EngineeringSenior-Site-Reliability-EngineerSenior-Site-Reliability-Engineering-ArchitectPrincipal-Site-Reliability-EngineerSenior-Reliability-EngineerSite-Reliability-Engineering-LeadSite-Reliability-Engineering-ManagerSenior
Job Description
About the RoleSunCore Digital is seeking a hands-on Senior Site Reliability Engineer to assess the current reliability and scalability of our systems, identify risks, and implement the technical changes required to address them.This is not a monitoring-only or advisory position. The SRE will investigate existing applications and infrastructure, establish reliability baselines, develop observability and testing capabilities, and directly implement reliability mitigations within the SRE domain. When a mitigation requires application-specific code or business-workflow changes, the system-owning team will implement and maintain those changes with SRE guidance.Because the platform has not yet been validated under real customer traffic, capacity and scalability will be treated as open risks until testing provides evidence otherwise.ResponsibilitiesReliability assessment and remediationAssess the reliability of web applications, Flutter mobile services, APIs, backend systems, infrastructure, databases, queues, and third-party integrations.Identify single points of failure, fragile dependencies, manual operational processes, and failure modes.Distinguish confirmed issues from suspected risks and areas that have not yet been evaluated.Design and implement shared reliability capabilities and improvements within the SRE domain.Define application-specific reliability changes and work with system-owning teams to implement them; those teams retain responsibility for their code and services.Transfer service-specific instrumentation, runbooks, and ongoing maintenance responsibilities to the appropriate system owners after the solution is tested and hardened.Track identified risks through implementation and validation rather than stopping at recommendations.Performance, load, and capacityEstablish a practical performance and capacity-testing program.Work with QA and product stakeholders to identify critical workflows and realistic usage scenarios.Establish baseline response times, throughput, concurrency, and resource consumption.Design and execute load, stress, endurance, scalability, and failure tests.Identify bottlenecks involving applications, databases, networks, queues, caches, infrastructure, and external services.Implement shared SRE mitigations and coordinate application-specific mitigation work with system-owning teams, which retain responsibility for their code and services.Repeat testing after changes to verify results.Document tested capacity, observed constraints, and remaining unknowns.Define safe operating limits and early-warning indicators.Performance and load testing have not yet been completed across the platform. Establishing this capability will be an early priority.ObservabilityAssess current logging, metrics, tracing, health checks, dashboards, and alerting.Establish consistent observability standards across services.Implement and harden shared health-checking and observability capabilities, transfer shared components to the designated long-term owner, and work with system-owning teams on service-specific instrumentation that they will maintain after handoff.Define meaningful service-level indicators and initial reliability objectives with system owners and engineering leadership.Build shared reliability dashboards and initial service-specific views, then train system owners to maintain their service-specific dashboards and alerts.Ensure alerts are actionable, routed to accountable owners, and tested.Identify monitoring blind spots.Improve application instrumentation in collaboration with developers.Ensure logs and telemetry do not expose sensitive information.Incident readinessHelp establish the initial on-call and escalation model.Create and test incident-response and troubleshooting runbooks.Define severity levels and technical escalation paths.Lead or support incident investigation.Improve detection, diagnosis, mitigation, and restoration capabilities.Facilitate technically focused post-incident reviews.Track corrective actions and recurring failure patterns.Conduct controlled failure exercises where appropriate.SRE automationAutomate repetitive reliability-engineering work, including health checks, diagnostics, alert enrichment, incident triage, capacity checks, and evidence collection.Develop safe automated remediation or self-healing for clearly defined and well-tested failure conditions.Reduce manual diagnostic, maintenance, incident-response, and reliability-validation steps within the SRE function.Define and implement the reliability checks, test logic, and SRE automation that should run through CI/CD. Work with DevOps to integrate them into the shared delivery framework, and with system-owning teams to maintain application-specific configuration after handoff.Create reusable reliability tools and patterns, harden and document them, transfer shared components to the designated long-term owner, and train application teams to operate the service-specific portions they own.Document automation
