Integrant is looking for game changers to join our team as " Lead AI Platform".The Lead AI Platform Engineer is responsible for bridging AI workloads with production-grade infrastructure, with a strong focus on NVIDIA AI stack, enabling high-performance, scalable, and optimized AI systems.This role focuses on model optimization, runtime efficiency, and GPU utilization, ensuring that AI workloads are production-ready, cost-efficient, and performant across enterprise environments.Roles and Responsibilities:Translate AI/ML workloads into optimized infrastructure and deployment strategiesOptimize model performance across GPU environments (latency, throughput, memory utilization)Design and implement inference and training pipelines using NVIDIA stack tools (TensorRT, Triton, NIM)Convert and optimize models across frameworks (PyTorch → ONNX → TensorRT)Analyze and resolve performance bottlenecks using profiling tools (GPU, memory, network)Improve GPU utilization and scheduling efficiency across clustersDesign scalable distributed training and inference architecturesWork closely with customers to define AI infrastructure strategies and deployment modelsSupport production deployments including monitoring, rollback, and performance validationConduct applied research to improve model efficiency and infrastructure utilizationMentor team members on AI infrastructure, optimization, and GPU systemsExperiment tracking tools (MLflow, W&B, Neptune) log parameters, metrics, and artifacts for comparisonFind the Model degradation happens post-deployment: concept drift, data pipeline changes, traffic pattern shiftsRoot cause analysis (RCA) applies to ML systems: isolating variables, reproducing issuesRequirements8+ years of experience in AI systems8+ years of experience in ML systems, HPC and AI infrastructureStrong proficiency in PythonStrong experience with GPU-based AI workloads and performance optimizationDeep understanding of model optimization techniques (quantization, pruning, batching)Hands-on experience with:PyTorchONNX / ONNX RuntimeTensorRT / TensorRT-LLMTriton Inference ServerKnowledge of CUDA, cuDNN, and GPU architecture fundamentalsExperience with distributed systems (multi-GPU / multi-node)Familiarity with:NCCL communicationNVLink / InfiniBandKubernetes or Slurm for orchestrationExperience deploying AI models into production environmentsAbility to analyze system bottlenecks (compute, memory, network)Experience with profiling tools (Nsight, TensorRT profiler, etc.)Knowledge of cost optimization strategies for GPU workloadsExperiment tracking tools (MLflow, W&B, Neptune) log parameters, metrics, and artifacts for comparisonFind the Model degradation happens post-deployment: concept drift, data pipeline changes, traffic pattern shiftsRoot cause analysis (RCA) applies to ML systems: isolating variables, reproducing issuesNice to HaveExperience with NVIDIA NIM and NGC ecosystemExposure to Megatron-LM, NeMo, or large-scale LLM training/inferenceExperience with LLM optimization techniques (KV cache, batching strategies)Familiarity with MLOps practices and CI/CD for AI systemsExperience in customer-facing architecture or consulting rolesFamiliarity with hybrid cloud / on-prem HPC environmentsBenefits Salary paid in USD Six-month career advancing opportunities Supportive and friendly work environment Premium medical insurance [employee +family] English language development courses Interest-free loans paid over 2.5 years Technical development courses Planned overtime program (POP) Employment referral program Premium location in Maadi Social insurance
Integrant is looking for game changers to join our team as " Lead AI Platform".
The Lead AI Platform Engineer is responsible for bridging AI workloads with production-grade infrastructure, with a strong focus on NVIDIA AI stack, enabling high-performance, scalable, and optimized AI systems.
This role focuses on model optimization, runtime efficiency, and GPU utilization, ensuring that AI workloads are production-ready, cost-efficient, and performant across enterprise environments.
Roles and Responsibilities:
- Translate AI/ML workloads into optimized infrastructure and deployment strategies
- Optimize model performance across GPU environments (latency, throughput, memory utilization)
- Design and implement inference and training pipelines using NVIDIA stack tools (TensorRT, Triton, NIM)
- Convert and optimize models across frameworks (PyTorch → ONNX → TensorRT)
- Analyze and resolve performance bottlenecks using profiling tools (GPU, memory, network)
- Improve GPU utilization and scheduling efficiency across clusters
- Design scalable distributed training and inference architectures
- Work closely with customers to define AI infrastructure strategies and deployment models
- Support production deployments including monitoring, rollback, and performance validation
- Conduct applied research to improve model efficiency and infrastructure utilization
- Mentor team members on AI infrastructure, optimization, and GPU systems
- Experiment tracking tools (MLflow, W&B, Neptune) log parameters, metrics, and artifacts for comparison
- Find the Model degradation happens post-deployment: concept drift, data pipeline changes, traffic pattern shifts
- Root cause analysis (RCA) applies to ML systems: isolating variables, reproducing issues
Requirements
- 8+ years of experience in AI systems
- 8+ years of experience in ML systems, HPC and AI infrastructure
- Strong proficiency in Python
- Strong experience with GPU-based AI workloads and performance optimization
- Deep understanding of model optimization techniques (quantization, pruning, batching)
- Hands-on experience with:
- PyTorch
- ONNX / ONNX Runtime
- TensorRT / TensorRT-LLM
- Triton Inference Server
- Knowledge of CUDA, cuDNN, and GPU architecture fundamentals
- Experience with distributed systems (multi-GPU / multi-node)
- Familiarity with:
- NCCL communication
- NVLink / InfiniBand
- Kubernetes or Slurm for orchestration
- Experience deploying AI models into production environments
- Ability to analyze system bottlenecks (compute, memory, network)
- Experience with profiling tools (Nsight, TensorRT profiler, etc.)
- Knowledge of cost optimization strategies for GPU workloads
- Experiment tracking tools (MLflow, W&B, Neptune) log parameters, metrics, and artifacts for comparison
- Find the Model degradation happens post-deployment: concept drift, data pipeline changes, traffic pattern shifts
- Root cause analysis (RCA) applies to ML systems: isolating variables, reproducing issues
Nice to Have
- Experience with NVIDIA NIM and NGC ecosystem
- Exposure to Megatron-LM, NeMo, or large-scale LLM training/inference
- Experience with LLM optimization techniques (KV cache, batching strategies)
- Familiarity with MLOps practices and CI/CD for AI systems
- Experience in customer-facing architecture or consulting roles
- Familiarity with hybrid cloud / on-prem HPC environments
Benefits
- Salary paid in USD
- Six-month career advancing opportunities
- Supportive and friendly work environment
- Premium medical insurance [employee +family]
- English language development courses
- Interest-free loans paid over 2.5 years
- Technical development courses
- Planned overtime program (POP)
- Employment referral program
- Premium location in Maadi
- Social insurance
Integrant, Inc. is a custom software development company focused on providing tailor made software solutions to fit your needs to a tee. We strive to uncover your pain points and identify how our team can seamlessly integrate with you and your business for a one-team approach. Our guiding principle is to always do the right thing for our customers and employees. Some days this means happy news of a “hit on the mark” demo, successful launch, or challenging problem solved.Other days this means making hard decisions, asking tough questions, or working more than we planned. Every day, it means doing our best to provide the highest quality service to each of our customers. We do that by investing our people in you and inspiring a people-to-people connection so when we say, “we share your goals,” we truly mean it. Contact us today to find out how we’re changing B2B. Website:https://www.integrant.comPhone number: +18587318700 Industry: Computer Software Company size: 51-200 employees Founded: 1992 Specialties:Software development, .NET, JavaScript, Java, iOS, Android, Xamarin, Selenium, Software Test Automation, Test Automation Framework, Agile Development, React Native, Ionic, Hybrid Mobile App Development, Mobile App Development, Web App Development, Desktop App Development, Software Quality Control, Custom Software Development, ReactJS, and AngularJS
Integrant, Inc. is a custom software development company focused on providing tailor made software solutions to fit your needs to a tee. We strive to uncover your pain points and identify how our team can seamlessly integrate with you and your business for a one-team approach.
Our guiding principle is to always do the right thing for our customers and employees.
Some days this means happy news of a “hit on the mark” demo, successful launch, or challenging problem solved.
Other days this means making hard decisions, asking tough questions, or working more than we planned.
Every day, it means doing our best to provide the highest quality service to each of our customers. We do that by investing our people in you and inspiring a people-to-people connection so when we say, “we share your goals,” we truly mean it.
Contact us today to find out how we’re changing B2B.
Website:https://www.integrant.com
Phone number: +18587318700
Industry: Computer Software
Company size: 51-200 employees
Founded: 1992
Specialties:
Software development, .NET, JavaScript, Java, iOS, Android, Xamarin, Selenium, Software Test Automation, Test Automation Framework, Agile Development, React Native, Ionic, Hybrid Mobile App Development, Mobile App Development, Web App Development, Desktop App Development, Software Quality Control, Custom Software Development, ReactJS, and AngularJS