About this role
Systems Engineer, HPC & Storage
Position Summary
Columbia University Irving Medical Center (CUIMC), part of a premier Ivy League institution and home to nearly 100 Nobel laureates, advances world-class biomedical and computational research spanning cancer genetics, genomics, molecular biology, and AI-driven discovery. The CUIMC Central IT (CUIMCIT) High-Performance Computing (HPC) team supports this mission through a rapidly evolving research computing environment featuring advanced GPU platforms, large-scale parallel storage, hybrid cloud integration, and next-generation scientific software.
We are seeking an experienced Systems Engineer, HPC & Storage to help modernize, operate, and support this mission-critical research infrastructure. The ideal candidate brings deep technical expertise, a strong service orientation, and the ability to collaborate effectively with faculty, postdoctoral researchers, graduate students, and technical teams. This position participates in a 24x7 operational model and may require occasional off-hours support for critical incidents.
Responsibilities
-
HPC Modernization & Infrastructure Engineering
-
Support the ongoing modernization of CUIMC’s HPC environment, including GPU-accelerated computing, high-speed interconnects, and next-generation clusters.
-
Contribute to lifecycle planning, hardware refreshes, and the integration of new compute and enterprise storage technologies.
-
Expand hybrid cloud capabilities and campus-wide research computing.
-
Server & Cluster Administration
-
Manage and maintain high-performance research computing servers, including Enterprise Linux distributions (Red Hat / Rocky Linux), SLURM/HPC clusters, and database infrastructures.
-
Monitor system performance, reliability, and capacity across compute nodes and shared infrastructure.
-
Enterprise Storage & Big Data Management
-
Administer large-scale parallel storage systems (such as VAST, Weka, or Isilon), including access control, quotas, advanced performance tuning, and backups.
-
Support the modernization of high-speed storage platforms to meet exponential data demands in genomics, high-resolution imaging, and AI workloads.
-
Scientific Software & Research Applications
-
Maintain, optimize, and deploy high-performance environments for scientific applications and tool chains such as MATLAB, Python, R, CUDA, PyTorch, Jupyter, container runtimes (Apptainer/Singularity), Grafana, Open On Demand.
-
Collaborate directly with researchers to troubleshoot computational workflows, optimize parallel jobs, and support emerging computational methodologies.
-
Virtualization & Cloud Integration
-
Manage KVM-based virtualization platforms and integrate on-premises infrastructure with campus and cloud resources.
-
Support containerization, and modern DevOps practices where applicable.
-
Identity & Access Management & Compliance
-
Administer user accounts, multi-protocol permissions, and group policies within enterprise Active Directory environments.
-
Enforce security best practices to ensure system stability, data integrity, and adherence to institutional compliance standards (e.g., HIPAA, NIST 800-53/171).
-
Resource Accounting & Collaboration
-
Manage computing resource allocations, usage tracking, and service metrics.
-
Represent CUIMCIT in technical committees, change-management processes, and cross-functional research computing initiatives.
-
Additional Duties
-
Perform other responsibilities as assigned by the HPC Team Director.
-
Minimum Qualifications
-
Bachelor’s degree or equivalent in education and experience, plus four (4) years of related experience
-
Expert-level proficiency in systems automation and scripting (e.g., Bash, Python, PowerShell).
-
Strong command of TCP/IP and IB networking, and virtualization technologies.
-
Excellent written/verbal communication skills and demonstrated customer-service orientation toward academic researchers.
-
Ability to lift equipment up to 50 lbs.
-
Preferred Qualifications
-
Minimum of 2–4 years of hands-on experience supporting large-scale technical environments, preferably in HPC or research computing.
-
Advanced degree in computer science, or related discipline.
-
Professional certifications in Linux administration, enterprise storage, networking, or cloud technologies.
-
Direct, hands-on experience managing enterprise parallel filesystems (e.g., VAST, Weka, Quobyte, Isilon) and workload managers (e.g., SLURM).
-
Background in academic medical centers or high-performance research computing environments supporting sensitive data (PHI).
-
Knowledge of modern ML/AI technologies.
-
Equal Opportunity Employer / Disability / Veteran
-
Columbia University is committed to the hiring of qualified local residents.