Principal Dev-Ops Architect
Cdit- Location
- Remote
- Workplace
- Remote
- Employment
- Full Time
- Salary
- —
Posted 4d ago
This is a remote position.
Senior technical authority for a cloud platform running global, multi-tenant SaaS services. This is a hands-on individual-contributor architect role — not people management. The architect designs the platform, sets standards and reference implementations other teams build on, and still writes Terraform, builds CI/CD pipelines, and stands up the AI/ML platform personally. Defines how reliability is measured against SLOs, how releases ship, and how the AI/ML platform is built and governed while keeping the environment HIPAA-compliant and SOC 2 Type 2 audit-ready. Influences products, software, and QA through architecture and example.
Core Responsibilities
- Own platform architecture and technical roadmap for infrastructure, deployment, observability, and the AI/ML platform.
- Set engineering standards, patterns, and golden paths for Infrastructure as Code (IaC), CI/CD, and AI tooling; drive adoption through reference implementations and architecture reviews.
- Manage all cloud infrastructure as code in Terraform — reusable modules, remote state, peer-reviewed PRs, drift detection, and automated plan/apply in CI/CD.
- Enforce policy-as-code (OPA, Sentinel, or equivalent) so infrastructure changes meet security and cost guardrails before merge.
- Design and deploy AWS infrastructure across dev, UAT, staging, and production for performance, availability, recoverability, and security (CIS Critical Security Controls).
- Build and operate CI/CD pipelines for large-scale applications on AWS; own release management, rollback, blue/green, canary, and release gates.
- Package and run containerized workloads on Docker and Kubernetes (EKS).
- Lead the SLI/SLO/SLA program and modern observability using OpenTelemetry; drive down MTTD and MTTR; lead blameless post-incident reviews and participate in on-call.
- Provision and operate the AI/ML platform — Anthropic Claude via AWS Bedrock and internal MCP services — all managed as IaC.
Build LLMOps practices
prompt versioning, evaluation pipelines, token cost attribution, guardrails, and audit logging of agent actions; enforce the PHI data boundary to BAA-covered providers only.
- Operate and evidence the platform controls required for SOC 2 Type 2 and HIPAA; own secrets management, supply-chain security (SBOM, image and dependency scanning), and FinOps.
Required Qualifications
- Bachelor's degree in Software Engineering or equivalent combination of technical education and work experience.
- 10+ years in SRE / DevOps / Platform Engineering delivering CI/CD, REST API deployment, containerization, IaaS/PaaS, data pipelines, and application observability — including time at a senior IC or architect level (Staff, Principal, or Architect).
Proven technical authority across teams
sets architecture and standards and influences delivery through expertise and example rather than direct management.
- Demonstrated experience driving adoption of a new practice or platform (IaC, CI/CD overhaul, or an AI/ML platform) across multiple teams.
- Hands-on Terraform, including reusable modules other teams consume via self-service, remote state, and change management in a CI/CD pipeline.
- Building and operating CI/CD pipelines for large-scale applications on AWS (GitHub Actions, Jenkins, GitLab, or AWS-native).
- Running containerized workloads on Docker and Kubernetes.
- Monitoring and troubleshooting using cloud-native tooling and OpenTelemetry.
- Linux system administration, Unix scripting, and automation.
- Experience working in a HIPAA / HITECH / HITRUST / PHI / PII or PCI DSS environment.
Skills
- Terraform
- HIPAA
- SOC 2
- AWS
- Docker
- Kubernetes
- EKS
- OpenTelemetry
- Anthropic Claude
- Bedrock
- Model Context Protocol
- LLMOps
- GitHub Actions
- Jenkins
- GitLab
- Linux
- Unix
- PCI DSS