Lab Administrator
ddn
Columbia Office
Posted Aug 20, 2026
- Full-time
- Department
Job description
**Role Summary** Responsible for the setup, maintenance, security, and day-to-day operation of a technical/research lab environment (servers, storage systems, networking equipment, and lab-issued workstations), ensuring high availability, performance, and compliance with organizational policies. **Key Responsibilities** - Install, configure, and maintain lab hardware (servers, storage arrays, NICs, switches) and software (OS images, drivers, monitoring tools) - Manage user access, accounts, and permissions across lab systems - Monitor system health, performance, and capacity (CPU, memory, storage, network utilization) - Troubleshoot hardware/software issues and coordinate with vendors for support/RMAs - Maintain documentation: network diagrams, asset inventories, configuration baselines, SOPs - Implement and enforce security policies (patching, firewall rules, access controls) - Manage backups, disaster recovery procedures, and data retention policies - Support researchers/engineers with environment setup for experiments, benchmarks, or testing (e.g., provisioning compute nodes, storage volumes, network configs) - Track licensing, warranties, and hardware lifecycle (procurement to decommissioning) - Coordinate lab scheduling/resource allocation if shared across teams **Required Skills/Qualifications** - Strong Linux administration experience (Ubuntu/RHEL/CentOS) - Networking fundamentals (TCP/IP, VLANs, bonding/LACP, basic troubleshooting) - Experience with storage systems (SAN/NAS, parallel filesystems like Lustre/GPFS a plus) - Scripting ability (Bash, Python) for automation - Familiarity with virtualization/containerization (KVM, Docker) is a plus - Experience with monitoring tools (Prometheus/Grafana, Nagios, Zabbix) - Understanding of hardware components (CPUs, NUMA architecture, PCIe, NICs/RDMA) - Good documentation and communication skills **Nice to Have** - Experience with HPC/AI infrastructure (InfiniBand, RDMA, GPU clusters) - Experience with configuration management (Ansible, Puppet)