Site Reliability Engineer I
American Express- Location
- Sunrise, FL, United States
- Workplace
- Hybrid
- Employment
- Full Time
- Salary
- USD 78,000–124,750/yr
Posted 19d ago
Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.
- Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms
- Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives
- Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics
- Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)
- Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure
- Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures
- Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability
- Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SRE‑owned work items
- Drive reliability improvements through sprint‑based commitments, including automation, operational fixes, and platform enhancements
- Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of business‑critical services
- Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements
- Support ITIL‑based Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and post‑implementation validation
- Monitor network performance and troubleshoot connectivity, latency, and access‑related issues impacting platform traffic
- Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates
- Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies
- Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability
- Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions
- Participate and lead the Development change review and change validation processes
- Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
- Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers
Education Qualifications
- Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.
- Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus
- Strong knowledge of operating systems and application runtimes such as Java and .NET
- Knowledge of distributed systems and service‑based architectures from an operations and reliability perspective
- Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana
- Knowledge of observability practices including logging, monitoring, tracing, and performance analysis
- Knowledge of RDBMS and NoSQL databases including MySQL, PostgreSQL, Couchbase, HBase, and Cassandra
- Knowledge of scripting and automation using languages such as PowerShell and Python
- Knowledge of AI, analytics, or AIOps platforms from an operational perspective is a plus
Work Experience
- Experience in Incident, Problem, and Change Management using ServiceNow or similar ITSM tools
- Experience supporting production systems in large‑scale enterprise environments with a focus on reliability and availability
- Experience in system administration, infrastructure operations, and network troubleshooting
- Experience with CI/CD pipeline implementation and support using tools such as Jenkins, GitHub Actions, XL Release (XLR), or similar
- Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms
- Knowledge of cloud‑based Site Reliability Engineering (SRE) practices with hands‑on experience on public cloud platforms such as AWS, Azure, or Google Cloud Platform
- Knowledge of containerization and orchestration technologies such as Docker and Kubernetes, and microservices‑based architectures
- Experience using enterprise monitoring and alerting platforms such as ELF
- Exposure to AI‑assisted monitoring, automation, or AIOps tools is a plus
- Proficiency in connecting to and administering servers via SSH (Secure Shell)
- Knowledge of core networking concepts including ports, protocols, firewalls, and secure remote access
Licenses & Certifications
- Certification in at least one programming language or runtime such as Java, .NET, or Python
- Certification in containerization and orchestration technologies (Docker, Kubernetes, OpenShift) is a plus
- Public cloud certification in AWS or GCP is a plus
- Certification or training related to AI platforms, analytics platforms, or AIOps is a plus
Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.
Skills
- TLS
- SOPS
- Java
- .NET
- Splunk
- Elasticsearch
- Prometheus
- Grafana
- MySQL
- PostgreSQL
- Couchbase
- HBase
- Cassandra
- PowerShell
- Python
- ServiceNow
- Jenkins
- GitHub Actions
- AWS
- Azure
- GCP
- Docker
- Kubernetes
- Shell
- OpenShift
- Express
More jobs at American Express
All 365Associate - Technology Vendor Management
American Express · Phoenix, AZ, United States · USD 78,000–124,750/yr · today
Senior Analyst/Analyst- Data Analytics
American Express · Gurugram, HR, India · today
Product Manager – IVA Experience Design
American Express · Phoenix, AZ, United States · Sunrise, FL, United States · New York, NY, United States · USD 103,750–174,750/yr · yesterday
Senior Software Engineer II - Full Stack Web Application
American Express · Chennai, TN, India · yesterday
Analyst-Data Science
American Express · Gurugram, HR, India · yesterday
Similar roles
EPIC Certified Systems Analyst (Beacon)
Inova · United States · USD 86,091–140,329 · today
EPIC Certified Systems Analyst Sr (Resolute HB/PB)
Inova · United States · USD 108,493–176,843 · today
Sr. Cloud Engineer I (6496)
MetroStar · Washington, DC · USD 128,000–151,000/yr · today
Field Service Technician, Battery Storage
Redwood Materials · McCarran, NV · today
Fusion Configuration Specialist – Oracle Cloud Financials
Mattel · El Segundo, CALIFORNIA, United States · USD 147,000–183,000/yr · today
IT Support Lead
Nadia Care · Philadelphia, Pennsylvania, United States · USD 70–80/hr · today