Chubb · Oracle Cloud · 2mo
SRE Lead
Malaysia- Posted
- 2026-07-14 (2mo)
- Place
- Malaysia
- Commitment
- Full Time
- Experience
- 5+ YOE
- Education
- Bachelor's
- Department
- Infrastructure Operations
- Source
- Oracle Cloud (the employer’s own listing)
Your match
Sign in to see how your skills match this job.
Skills in this posting
PythonJavaSDLCDockerKubernetesCI/CDTerraformPrometheusGrafanaSplunkObservabilitySite Reliability Engineering
Key Objective:
• Lead the Site Reliability Engineering function to define and drive the organisation’s reliability engineering strategy — bridging software development and operations through engineering discipline, not manual process.
• Own the end-to-end reliability posture of production systems: define SLO/SLI frameworks, govern error budgets, and enforce production-readiness standards to protect business continuity.
• Build, mentor, and scale a high-performing SRE team that prioritises engineering over toil — automating manual work, embedding reliability into the SDLC, and driving down mean time to recovery through systematic improvement.
• Champion observability-led engineering through full-stack Dynatrace adoption, AIOps integration, and data-driven reliability decision-making at every layer of the stack.
• Serve as the primary reliability engineering partner to development and platform leadership, shaping architecture decisions, release policies, and automation strategy.
Key Responsibilities:
• Define and drive the SRE strategy and multi-year roadmap aligned to business priorities.
• Lead and develop the SRE team, including hiring, onboarding, performance management, career development, and succession planning.
• Own incident management, including severity classification, escalation, response SLAs, and leadership of major incidents.
• Champion blameless postmortems, root cause analysis, and implementation of systemic fixes.
• Establish and govern SLOs, SLIs, and error budgets, ensuring reliability targets are aligned to business needs.
• Drive resilience engineering, including chaos engineering, GameDays, production readiness reviews, and failure mode analysis.
• Reduce toil through automation, improved runbooks, and continuous operational improvement.
• Own observability and alerting standards, including monitoring strategy, dashboards, and alert quality.
• Partner with engineering, architecture, product, and leadership teams to embed reliability into design and delivery.
• Represent the SRE function in senior forums and provide reporting on reliability, risk, and operational performance.
Qualifications:
• Degree in Computer Science, Software Engineering, IT, or a related technical field.
• 8+ years’ experience in software engineering, platform reliability, or SRE, including 3+ years in people leadership.
• Strong hands-on coding ability in Python, Go, or similar, with experience building automation and self-healing solutions.
• Proven experience leading enterprise-scale SRE or platform reliability functions.
• Experience defining and operating RTO/RPO and SLI/SLO frameworks.
• Strong background in observability, production readiness, error budgets, and chaos engineering.
• Experience leading on-call models, incident response, and executive stakeholder engagement.
• Solid understanding of SDLC, Agile, and DevOps delivery models.
• ITIL Foundation is desirable.
Managerial & Soft Skills:
• Proven people leader with experience building and coaching high-performing teams.
• Strategic thinker who can turn business priorities into reliability roadmaps.
• Strong communicator who can explain technical risk in business terms.
• Calm and decisive during major incidents.
• Influential partner across engineering, product, and leadership teams.
• Strong advocate for developer experience and sustainable on-call practices.
• Data-driven and able to balance reliability, speed, and cost.
• Champions psychological safety, continuous learning, and operational excellence.
Technical Skills:
• Expert in Dynatrace, with experience in observability, monitoring, SLOs, tracing, and log management.
• Proficient in Grafana, Prometheus, Splunk, ELK, Azure Monitor, and Log Analytics.
• Strong knowledge of OpenTelemetry and telemetry pipeline design.
• Experience with ServiceNow, CI/CD tools, Kubernetes, Docker, Terraform, and Bicep.
• Familiar with Java, .NET, databases, APIs, Kafka, and cloud platforms, especially Azure.
• Experience with AIOps, AI-assisted triage, and automation tooling.
• Able to support reliability engineering through scripting, auto-remediation, and operational automation.
Desired:
• Experience in insurance or financial services.
• Dynatrace, Azure, ITIL 4, or Google Cloud/SRE-related certifications.
• Experience with chaos engineering, AIOps, MLOps, and FinOps.
• Strong analytical skills and experience working with large operational datasets.
Chubb
12 open roles in Malaysia, straight from Chubb’s own careers page.
- Assistant Manager, SMEMalaysia · 1w
- APAC End User Services LeadMalaysia · 2w
- Senior Data AnalystMalaysia · 1mo
- Business Analyst, P&C ProductsMalaysia · 1mo
- Customer Fulfillment SpecialistMalaysia · 1mo
- Consumer Servicing Delivery LeadMalaysia · 1mo
- Marine UnderwriterMalaysia · 2mo
- Executive, Credit ControllerMalaysia · 2mo
- Senior Underwriting Executive, A&HMalaysia · 2mo
- Business Analyst, Project ManagementMalaysia · 2mo
- APAC HR Services Hub AssociateMalaysia · 2mo