Senior Site Reliability Engineer
Oracle- Location
- BENGALURU, KARNATAKA, India
- Workplace
- —
- Employment
- —
- Salary
- —
Posted 8d ago
Job Responsibilities
- Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
- Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
- Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
- Build automation and tooling to reduce operational toil and improve production safety.
- Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
- Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
- Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
- Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
- Contribute to incident-management practices, operational readiness, and service ownership improvements.
- Share technical knowledge and support team members through documentation, reviews, and collaboration.
- Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.
Mandatory Skills
- 4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
- Experience operating and improving highly available production systems.
- Strong programming or scripting skills in Python, Java, Go, or similar languages.
- Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
- Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
- Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
- Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
- Experience with deployment pipelines, release validation, automation, and change-management practices.
- Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
- Ability to work independently on technical problems and collaborate effectively with engineering teams.
- Strong written and verbal communication skills.
Preferred Skills
- Experience with OCI and cloud infrastructure services.
- Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
- Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
- Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts.
- Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
- Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
- Familiarity with security, compliance, and access-control practices in production environments.
Self-Test Questions
- Do you have 4–8 years of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
- Have you independently operated or improved a production service, system, or infrastructure component?
- Can you investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions?
- Do you have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
- Are you proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
- Have you built or improved automation, deployment validation, CI/CD pipelines, or operational tooling?
- Do you have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs?
- Can you work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation?
Career Level - IC3
Skills
- OCI
- Python
- Java
- Go
- Linux
- Kubernetes
More jobs at Oracle
All 376Senior Platform Software Engineer
Oracle · Nashville, TN, United States · United States · USD 92,500–209,500/yr · today
Senior Platform Software Engineer - OCI
Oracle · Austin, TX, United States · Nashville, TN, United States · United States · USD 92,500–209,500/yr · today
Lead Principal Product Manager
Oracle · Nashville, TN, United States · Seattle, WA, United States · USD 119,200–264,100/yr · today
Data Center Build Engineer
Oracle · BANGKOK, Thailand · yesterday
Principal Software Engineer, Core Infrastructure
Oracle · Nashville, TN, United States · USD 114,600–234,600/yr · yesterday
Similar roles
Senior DevOps Engineer
Bluehost · Mumbai, India · today
Technical Support Engineer 2
Twilio · India · today
Senior Network Engineer
Twilio · India · today
Sr. Staff Site Reliability Engineer
Zscaler · Bangalore, IND · today
Staff Software Development Engineer - DevOps
Zscaler · Bangalore, IND · today
Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)
Zscaler · Bangalore, IND · today