Senior Infrastructure Engineer
jobgether
India
Posted Sep 8, 2026
- Full-time
- Remote
- Security & IT
Job description
**This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Infrastructure Engineer based in Canada.** As a Senior Infrastructure Engineer, you will take ownership of business-critical infrastructure supporting demanding AI and GPU workloads. You will design, deploy, and operate large-scale OpenStack and Kubernetes environments with a strong focus on performance, reliability, and scalability. The role combines hands-on systems engineering with automation, infrastructure-as-code, observability, security, and incident response. You will work directly with physical servers, racks, networking, storage, and GPU infrastructure across data-centre environments. Your work will have a visible impact on platform stability, customer experience, and the ability to scale globally. You will collaborate closely with platform, DevOps, AI, product, and support teams in a fast-moving, technically ambitious environment. This is an ownership-driven opportunity for an engineer who enjoys solving complex infrastructure problems and improving systems end-to-end. ### Accountabilities: - Design, deploy, operate, and continuously improve OpenStack and Kubernetes environments supporting high-performance GPU workloads. - Take end-to-end ownership of infrastructure performance, scalability, resilience, and operational reliability. - Build and maintain infrastructure using infrastructure-as-code, GitOps, CI/CD, and Git-based operational practices. - Automate provisioning, deployment, configuration, and recurring infrastructure workflows to improve efficiency and consistency. - Optimize GPU workload scheduling using Kubernetes and NVIDIA technologies to maximize platform performance and resource utilization. - Implement and improve monitoring, logging, alerting, and observability systems to identify and resolve infrastructure issues proactively. - Lead incident response activities, troubleshoot complex production problems, and drive post-incident improvements that strengthen overall reliability. - Maintain robust security controls across infrastructure and container environments, including RBAC, network policies, access controls, and tenant isolation. - Build, rack, cable, configure, and commission physical server and GPU infrastructure within data-centre environments. - Work with networking and storage systems to ensure reliable, scalable infrastructure for compute-intensive workloads. - Collaborate with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with technical and customer requirements. - Contribute to the continuous evolution of infrastructure architecture and operational practices as global environments scale. - Identify opportunities to improve automation, performance, reliability, and operational efficiency across the infrastructure platform. ## Requirements - Extensive hands-on experience administering Linux systems, with strong technical depth across operating systems, troubleshooting, and infrastructure operations. - Proven experience physically building, assembling, cabling, configuring, and commissioning servers and racks. - Direct hands-on experience working in data centres, including physically installing and racking hardware on-site. - Willingness and ability to travel to data-centre sites in Quebec as required. - Strong understanding of networking and storage technologies and their role within large-scale infrastructure environments. - Experience installing, racking, and configuring GPU hardware is highly advantageous, particularly with NVIDIA platforms. - Production experience operating OpenStack and/or Kubernetes at scale is strongly preferred. - Experience with infrastructure automation, infrastructure-as-code, CI/CD pipelines, GitOps, and Git-based workflows is an advantage. - Exposure to high-performance computing, large-scale compute, or other demanding infrastructure environments is desirable. - Experience with Kubernetes GPU scheduling and NVIDIA tooling is a plus. - Strong troubleshooting and analytical abilities, with a methodical approach to diagnosing complex infrastructure issues. - Ability to take ownership of systems from design through deployment, operation, troubleshooting, and continuous improvement. - Comfortable working in a fast-paced environment where priorities can evolve and engineers are trusted to make decisions independently. - Strong collaboration and communication skills, with the ability to work effectively across technical and non-technical teams. - Contributions to open-source infrastructure or technology projects are a plus. ## Benefits - Competitive salary. - Annual discretionary bonus scheme. - Employee wellbeing benefits. - 25 days of annual holiday plus public holidays. - Flexible working arrangements, with remote or hybrid options depending on role and location. - Significant autonomy and ownership, with the freedom to take initiative and experiment. - Opportunity to work on challenging, high-performance AI and GPU infrastructure. - Visible impact on infrastructure reliability, scalability, and customer experience. - Clear career progression and professional growth opportunities. - Collaborative international environment built around trust, transparency, and ownership. - Opportunity to contribute to the evolution of a rapidly scaling AI infrastructure platform and its engineering culture.