Software Engr II
Honeywell- Location
- Bengaluru, Karnataka, India
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 1mo ago
Job Title
Senior Site Reliability Engineer
Location
Bangalore
We are seeking a highly technical SRE Engineer to design, build, and maintain fault-tolerant, scalable, and highly available distributed systems. You will champion SRE best practices, reduce manual operations (toil) via automation, and partner with product development squads to embed reliability into the software delivery lifecycle
Your role will include overseeing, supervising and reviewing tasks performed by team members to ensure effective execution of work; managing end-to-end processes and projects for both internal and external clients with responsibility for timely and accurate delivery; issuing clear instructions and directions to team members on tasks to be performed; and mentoring and guiding junior colleagues to Support their skill development, professional growth, and overall success.
Key Responsibilities
Reliability & Availability
- Ensure high availability and uptime of production services.
- Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs).
- Design and implement disaster recovery (DR) strategies, including RTO/RPO targets.
- Conduct capacity planning and scalability assessments.
Monitoring & Observability
- Implement and maintain monitoring, logging, and alerting systems.
- Create dashboards to track system health and performance.
- Improve observability using tools such as Prometheus, Grafana, Azure Monitor, Dynatrace, or Elastic.
- Proactively detect, investigate, and resolve system issues.
- Incident Management
- Participate in on-call rotations and incident response.
- Lead troubleshooting during service disruptions.
- Perform Root Cause Analysis (RCA) and drive corrective actions.
- Reduce Mean Time To Detect (MTTD) and Mean Time To Recover (MTTR).
Automation & Engineering
- Develop automation to eliminate repetitive operational tasks.
- Build self-healing and auto-scaling capabilities.
- Create and maintain Infrastructure as Code (IaC).
- Improve CI/CD pipelines and deployment reliability.
Cloud & Infrastructure Management
- Manage cloud environments and production infrastructure.
- Support Kubernetes/AKS clusters, databases, networking, storage, and messaging services.
- Optimize infrastructure costs and resource utilization.
- Ensure security, compliance, and operational best practices.
Collaboration
- Partner with software development teams throughout the software lifecycle.
- Participate in architecture reviews and non-functional requirement (NFR) assessments.
- Review reliability risks and recommend improvements.
- Support release planning and production readiness reviews.
Required Technical Skills
Infrastructure & Cloud
- Azure/AWS or GCP
- Kubernetes / AKS / OpenShift
- Linux administration
- Networking fundamentals (TCP/IP, DNS, Load Balancing)
- Database administration basics (SQL/NoSQL)
- Automation & DevOps
- Terraform, ARM, Bicep, or CloudFormation
- CI/CD tools (Azure DevOps, GitHub Actions, Jenkins, GitLab)
- Infrastructure as Code (IaC)
- Configuration management tools such as Ansible
- Programming such as Python, Jave, Sheel scripting, Go
Monitoring & Observability
- Grafana
- Prometheus
- Dynatrace
- Elastic Stack
- Azure Monitor / Log Analytics
Experience Level
5+ yrs
Required Qualifications
- Education: Bachelor's/Master’s degree in Computer Science, Information Technology, or equivalent practical experience.
- Experience supporting large-scale cloud-native applications.
- Experience with incident management and production support.
- Understanding of SRE concepts such as:
- Error Budgets
- SLI/SLO/SLA
- Chaos Engineering
- High Availability
- Reliability Engineering
Success Metrics
- Service Availability (% Uptime)
- SLA/SLO Compliance
- MTTR / MTTD
- Deployment Success Rate
- Production Incident Reduction
- Operational Cost Optimization
- Automation Coverage
- Customer Experience Metrics
Skills
- Prometheus
- Grafana
- Azure Monitor
- Dynatrace
- Elastic
- Kubernetes
- AKS
- Azure
- AWS
- GCP
- OpenShift
- Linux
- TCP/IP
- DNS
- SQL
- Terraform
- ARM
- Bicep
- AWS CloudFormation
- Azure DevOps
- GitHub Actions
- Jenkins
- GitLab
- Ansible
- Python
- Go
- ELK Stack
- Azure Log Analytics
More jobs at Honeywell
All 302Sr Field Service Technician | HVAC Controls Building Automation
Honeywell · Acton, MA, United States · USD 67,000–83,000/yr · today
Lead Field Service Technician | Building Automation HVAC Controls
Honeywell · Acton, MA, United States · USD 69,000–87,000/yr · today
Lead Field Service Technician | Fire Alarm Systems
Honeywell · Acton, MA, United States · USD 69,000–87,000/yr · today
Lead Field Service Technician
Honeywell · Des Moines, IA, United States · Omaha, NE, United States · Iowa City, IA, United States +1 · today
Advanced Software Engineer- Full Stack Cloud Developer
Honeywell · Atlanta, GA, United States · today
Similar roles
Staff Product Security Engineer, PSIRT
ServiceNow · Hyderabad, Telangana, India · today
Sr Staff Inbound Product Manager
ServiceNow · Hyderabad, Telangana, India · today
Sr Software Engineer - Cloud Platform—Kubernetes Development
ServiceNow · Hyderabad, India · today
Senior Software Engineer
Smiths Group · Bengaluru, KA, India · today
QA Engineer Contractual - 6 months
AbhiBus · Hyderabad, TS, India · yesterday
Sr Staff Software Engineer
ServiceNow · Hyderabad, Telangana, India · today