LLM Inference & GPU Systems Consultant
Delan Associates- Location
- Charlotte, North Carolina, United States
- Workplace
- —
- Employment
- Contract
- Salary
- —
Posted 1mo ago
Job Title: LLM Inference & GPU Systems Consultant
Location: Charlotte, NC (Onsite)
Duration: 6+ Months
Must be onsite at client in Charlotte, NC at least 3 days/week
Role Overview
We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.
Key Responsibilities
NVIDIA GPU Runtime Optimization
Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.
Inference Serving
Deploy and manage inference engines including vLLM and TensorRT-LLM.
Hardware Utilization
Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.
Model Lifecycle Management
Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.
Platform Operations
Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.
Required Qualifications
8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).
Proficiency in OpenShift AI and GPU orchestration tools like RunAI.
Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.
Proven track record managing the Hugging Face deployment lifecycle.
Skills
- LLM
- Generative AI
- OpenShift
- Llama
- vLLM
- TensorRT
- Kubernetes
- Hugging Face
More jobs at Delan Associates
All 260Application Developer_3 (109)
Delan Associates · Brooklyn, New York, United States · yesterday
Application Developer_2 (108)
Delan Associates · Brooklyn, New York, United States · yesterday
Application Developer_1 (107)
Delan Associates · Brooklyn, New York, United States · yesterday
Programmer Analyst (106)
Delan Associates · Brooklyn, New York, United States · yesterday
Senior IT Specialist (104)
Delan Associates · Brooklyn, New York, United States · yesterday
Similar roles
Software Engineer II, Data Analytics & Engineering
Pinterest · San Francisco, CA, US · Remote, US · today
Software Engineer, Agent Delivery
Pallet · San Francisco · USD 135,000–180,000/yr · today
Manager II, Software Engineering, Infrastructure
Samsara · SF Bay Area · USD 154,700–260,000/yr · today
Intern, Software Engineer AI Agents (Fall 2026/Winter 2027)
Bot Auto · Houston, TX · today
Staff Software Engineer, Backend
Archer · San Jose, California, United States · today
Software Engineer, Sandbox & Agent Executor
Retool · San Francisco · USD 401–163,800/yr · today