Principal Data Center Infrastructure Software Engineer
designworkstalent
Bellevue
Posted Aug 14, 2026
- Full-time
- Remote
- Engineering
Job description
# Data Center Infrastructure Software Engineer **Location:** Hybrid | Bellevue, WA Area **Titles:** Senior | Staff | Principal (multiple roles available) # Build the Data Center Software Infrastructure ## **About the Opportunity** A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications. Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads. We're seeking **Data Center Software Engineers** to lead the design, development, configuration, and automation of AI infrastructure clusters. ## **The Opportunity** Your responsibility begins once servers and racks are installed in the data center and extends through software deployment, networking, configuration, cluster bring-up, and automation, ensuring the platform is fully operational and ready for customer workloads. ## **What You'll Do** - Develop infrastructure-as-code, automation, and provisioning systems for compute, networking, and storage. - Deploy and optimize Kubernetes, container, and distributed computing platforms. - Optimize GPU, networking, storage, and system performance for large-scale AI workloads. - Troubleshoot complex issues across hardware, operating systems, networking, storage, and software stacks. - Build reliability, observability, and operational excellence practices for mission-critical infrastructure. ## **What We're Looking For** - 5+ years of experience designing, building, or operating large-scale Linux-based infrastructure. - Hands-on experience with Kubernetes, containerization, and distributed systems in production environments. - Experience with infrastructure-as-code and automation tools such as Terraform, Ansible, or similar frameworks. - Strong experience operating cloud or datacenter-scale infrastructure. ## **Preferred Qualifications** - Experience with bare-metal provisioning and hardware lifecycle management platforms (e.g., MAAS, Ironic, xCAT, xCAT, or similar). - Experience with IPMI, Redfish, PXE boot, and automated operating system deployment at scale. - Experience managing GPU clusters in datacenter or cloud environments. ## **Compensation** - Competitive base pay for Bellevue market - Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance - U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays. ## **Location** - Hybrid role based in the Bellevue, WA area. - Approximately three days per week in the office. - Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply. - U.S. work authorization is required. Visa sponsorship is not currently available. ## **Why Join?** - Ground-floor opportunity: you'll be among the earliest engineers on the team, directly shaping architecture, tooling, and culture. - Work directly on cutting-edge AI infrastructure at real scale — from data center design through GPU clusters to production inference.