AI Solution Architect
Uvation- Location
- India
- Workplace
- Remote
- Employment
- Full Time
- Salary
- —
Posted today
Job Overview
We are seeking an experienced AI Solution Architect to design and lead end-to-end enterprise AI Factory and GPU infrastructure solutions spanning compute, high-performance networking, storage, Kubernetes, cloud, and AI/ML platforms. The role requires strong expertise in NVIDIA GPU technologies, AI workloads, scalable infrastructure architecture, security, observability, performance engineering, and capacity planning.
Key Responsibilities
- Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
- Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing, and high-performance computing.
- Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch, and GPU resource allocation.
- Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.
- Design AI storage and data architectures using object storage, parallel file systems like Ceph, WEKA, , or equivalent platforms.
- Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.
- Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.
- Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations, and implementation roadmaps.
- Lead technical evaluations, proof-of-concepts, vendor assessments, and architecture review boards.
- Collaborate with infrastructure, network, security, storage, cloud, data, application, and operations teams.
- Define performance, availability, scalability, security, and cost objectives and validate architecture against measurable acceptance criteria.
- Provide technical leadership during deployment, migration, integration, troubleshooting, and production transition.
- Required Technical Skills
AI / ML Architecture
- NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystem.
- PyTorch, TensorFlow, JAX and operational understanding of training and inference workloads.
- GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.
- LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.
GPU & AI Factory Infrastructure
- NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; familiarity with next-generation systems.
- NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.
- DGX/HGX/OEM GPU server architecture and lifecycle management.
- AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.
High-Performance Networking
- 100/200/400/800G Ethernet, InfiniBand, RoCEv2 and RDMA, Netris
- NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.
- BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, QoS and congestion management.
- GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.
AI Storage & Data Architecture
- Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.
- Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.
- Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.
- GPUDirect Storage and storage/network performance optimization.
AI Platform & Orchestration
- Kubernetes, GPU Operator, container runtimes and Kubernetes GPU scheduling.
- HPC or other equivalent workload schedulers.
- Model serving/inference platforms and MLOps platform architecture.
- API gateways, service discovery, secrets management and platform integration.
Cloud & Hybrid Architecture
- AWS and/or Azure AI infrastructure and security services.
- Hybrid cloud connectivity, IAM, private networking, cloud storage and workload placement.
- Cloud cost optimization, capacity planning and FinOps considerations for GPU workloads.
Security & Governance
- Zero Trust, network segmentation, IAM/RBAC, PAM and workload identity.
- GPU, DPU, container, Kubernetes, firmware and supply-chain security.
- Encryption at rest/in transit, secrets management, audit logging and compliance controls.
- AI-specific risks including data/model protection, tenant isolation and secure model access.
Observability & Reliability
- Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.
- Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.
- High availability, backup/restore, disaster recovery, business continuity and failure-domain design.
- Performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.
Architecture Deliverables
- AI Factory reference architecture and solution blueprints
- High-Level Design (HLD) and Low-Level Design (LLD)
- Network, compute, GPU and storage architecture diagrams
- Capacity, performance and scalability models
- Technology evaluation and vendor comparison documents
- Security architecture and threat-model inputs
- Bill of Materials (BOM) and infrastructure sizing
- Migration/deployment strategy and implementation roadmap
- Operational readiness checklist, runbooks and acceptance criteria
Experience & Qualifications
- 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.
- Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.
- Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.
- Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.
Preferred Certifications
- NVIDIA certifications or equivalent GPU/AI infrastructure credentials
- AWS Solutions Architect / Azure Solutions Architect
- TOGAF or equivalent enterprise architecture certification
- CCNP/CCIE or equivalent networking certification
- CISSP or equivalent security certification
- Kubernetes certifications such as CKA/CKAD
- Red Hat / Linux certifications
Skills
- Kubernetes
- Ceph
- Machine Learning
- CUDA
- PyTorch
- TensorFlow
- JAX
- LLM
- Generative AI
- Retrieval-Augmented Generation
- RDMA
- BGP
- MLOps
- AWS
- Azure AI
- IAM
- Google Cloud Storage
- Zero Trust
- RBAC
- Prometheus
- Grafana
- OpenTelemetry
- Linux
- Azure
- CISSP
More jobs at Uvation
All 67Similar roles
Global Partner Solution Engineer – Cisco & Splunk
Cis · Bangalore, India · just now
DE&A - Core - Snowflake Solution Architect
Zensar Technologies · Bangalore, Karnataka, India · today
AI Solution Architect - MQ DIA
528 Eli Lilly Asia Pacific SSC Sdn Bhd · India, Bangalore · today
AI Solutions Engineer
Hp · Bengaluru, Karnataka, India · today
AI Solutions Engineer
Hp · Bengaluru, Karnataka, India · today
Associate (AI) Solution Consultant - Orbit Program
Celonis · Bangalore, India · yesterday