ML Infrastructure Engineer
Clera- Location
- San Mateo
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 9d ago
About the Role
This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context and data governance layer for AI agents in highly regulated industries. You will own the inference and model-serving infrastructure end to end, ensuring AI agents run reliably, accurately, and at scale in production environments where performance is non-negotiable.
What You'll Do
- Design, build, and operate inference and model-serving infrastructure from development through production deployment.
- Scale systems to support AI agents running reliably under increasing concurrency and production load.
- Identify and resolve infrastructure bottlenecks in close collaboration with ML and platform engineering teams.
- Optimize systems for latency, throughput, and reliability at scale.
What We're Looking For
- 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.
- Hands-on experience designing and scaling inference serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.
- Strong systems engineering fundamentals with expertise in distributed systems, containerization, and orchestration (Docker, Kubernetes).
- Demonstrated ability to optimize production ML systems for latency, throughput, and reliability under high concurrency.
- Experience with cloud infrastructure platforms such as AWS, GCP, or Azure for deploying and managing ML workloads.
- Proficiency with monitoring, observability, and debugging tools such as Prometheus, Grafana, ELK, or distributed tracing frameworks.
- Proficiency in at least one systems programming or backend language: Python, Go, Rust, C++, or Java.
- Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune) is a plus.
- Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines is a plus.
- Experience with enterprise data infrastructure, data pipelines, or data integration platforms is a plus.
Location
This role is on-site in San Mateo, California. Visa sponsorship is not available.
Skills
- Machine Learning
- TensorFlow
- Triton
- Docker
- Kubernetes
- AWS
- GCP
- Azure
- Prometheus
- Grafana
- ELK Stack
- Python
- Go
- Rust
- C++
- Java
- Neo4j
- Amazon Neptune
More jobs at Clera
All 258Similar roles
IT Support Lead
Nadia Care · Philadelphia, Pennsylvania, United States · USD 70–80/hr · today
Network Engineer
Delart · Austin, TX · today
Computing Undergraduate Student Intern: DevOps Internship Program - Summer 2027
Llnl · Livermore, CA, United States · USD 23–32/hr · today
IT Support Assistant (Onsite - Phoenix, AZ)
Intelligent Technical Solutions · Phoenix, AZ · USD 15–16/hr · today
Senior Cloud Infrastructure Engineer
Anduril Industries · Washington · District of Columbia · United States · USD 146,000–194,000/yr · today
Senior AI Infrastructure Engineer, Physical Infrastructure
Anduril Industries · Costa Mesa, California, United States · USD 166,000–220,000/yr · today