Engineering Manager, SRE
jobgether
Switzerland
Posted Sep 11, 2026
- Full-time
- Remote
- Security & IT
Job description
**This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engineering Manager, SRE based in Switzerland.** This is a hands-on engineering leadership role responsible for building a highly reliable foundation for a globally distributed technology platform. You will lead a Site Reliability Engineering team while remaining deeply involved in technical direction and complex infrastructure challenges. The role combines people leadership with expertise across Kubernetes, AWS, PostgreSQL, CI/CD, observability, infrastructure as code, and reliability engineering. You will shape how the team balances operational excellence, incident response, reliability improvements, and longer-term engineering initiatives. A key focus will be maturing SLOs, error budgets, observability, and reliability practices across the wider engineering organization. You will also act as a trusted technical and organizational partner to engineering, security, and senior leadership. The environment is fully remote and asynchronous, offering significant autonomy in a fast-growing, globally distributed organization. ### Accountabilities - Lead and develop a Site Reliability Engineering team, owning the full career lifecycle of direct reports including onboarding, feedback, performance management, progression, coaching, and hiring. - Establish a clear team direction and priorities aligned with broader company goals, balancing operational commitments with project delivery and protecting the team’s focus. - Serve as the team's spokesperson across engineering and with senior leadership, communicating priorities, progress, risks, and technical challenges clearly. - Own SRE delivery goals, deciding what the team commits to, how work is prioritized, and how operational responsibilities are managed. - Design and maintain effective support rotations and on-call processes while strengthening incident response practices. - Provide technical leadership across Kubernetes, AWS, PostgreSQL, DNS and TLS, CI infrastructure, and the broader infrastructure platform. - Guide the development of reliability practices including SLOs, error budgets, observability, incident response, and post-incident improvements. - Partner closely with Security on infrastructure threats, patching, controls, audits, and compliance obligations. - Manage relationships with infrastructure and platform vendors, including renewals and commercial discussions with support from senior leadership. - Remain hands-on enough to review technical work, challenge architectural decisions, participate credibly in incidents, and identify emerging reliability issues before they escalate. - Build strong relationships across engineering and encourage teams to bring operational and reliability challenges forward early. - Continuously improve team health, collaboration, conflict resolution, and retrospective practices. ### **Requirements:** - Proven experience leading an SRE, infrastructure, platform engineering, DevOps, or similarly focused technical team, with direct responsibility for performance and career development. - Strong hands-on background in site reliability, DevOps, or cloud infrastructure engineering, with sufficient technical depth to review designs, challenge implementation decisions, and contribute during production incidents. - Production experience with Kubernetes and AWS at meaningful scale, including the operational realities of running cloud infrastructure. - Hands-on experience building, enabling, or scaling AI infrastructure and working with AI-related engineering workloads. - Strong understanding of observability principles and practices, infrastructure as code with Terraform, and CI/CD platforms such as GitLab CI, GitHub Actions, or Jenkins. - Experience with Docker, shell scripting, and production infrastructure operations. - Proven ownership of reliability practices including incident response, on-call operations, SLOs, error budgets, and turning incidents into lasting engineering improvements. - Experience working in regulated environments, with an understanding of infrastructure controls, compliance, and security requirements. - Exceptional prioritization skills, particularly when operational workloads compete with project commitments. - Excellent written communication and documentation skills, with the ability to lead effectively in a highly distributed and asynchronous environment. - Strong relationship-building, collaboration, conflict-resolution, and stakeholder-management capabilities. - A coaching-oriented leadership style, with evidence of developing engineers both technically and professionally. - Strong judgment, accountability, adaptability, curiosity, and commitment to high-quality execution. - Nice-to-have experience with Elixir, Java, Clojure, Node.js, Python, or another backend programming language. - Additional desirable experience includes OpenTelemetry, distributed tracing, Honeycomb, PostgreSQL or Aurora performance optimization, connection pool management, query tuning, Linux systems administration, security, FinOps, and cloud cost management. - Experience growing an engineering team from a small base and establishing a strong hiring bar is advantageous. - Ability to work effectively across global teams and time zones. ### **Benefits:** - Annual salary range of **USD $75,450–$169,700**, with actual compensation determined by location, experience, skills, training, business needs, and market conditions. - Fully remote, work-from-anywhere environment. - Flexible working hours within an asynchronous work culture. - Flexible paid time off. - **16 weeks of paid parental leave**. - Budget for coworking spaces, learning, and wellness activities, including gym memberships. - Mental health support services. - Stock options. - Home office budget and IT equipment. - Global exposure through collaboration with colleagues across multiple continents. - Opportunities to travel internationally and meet colleagues at company events. - A high-autonomy environment where employees are encouraged to organize their schedules around their lives and personal commitments. - Opportunity to influence the maturity of reliability engineering practices while working on complex infrastructure and platform challenges.