Roles & Responsibilities¶
This section outlines my key roles and responsibilities in supporting large-scale Linux and high-performance computing (HPC) environments in production settings.
HPC Operations & Administration¶
- Manage and support multi-petaflop HPC environments for research workloads.
- Perform tasks like setting up and managing compute nodes, from installation to decommissioning.
- Support workload scheduling using PBS Pro to ensure tasks run smoothly in shared environments.
- Ensure stable job execution across both CPU and GPU-based clusters.
System Monitoring & Reliability¶
- Monitor system health and availability using tools like NVIDIA Bright Cluster Manager, Dell OpenManage Enterprise, and InfiniBand UFM.
- Run proactive health checks to identify potential issues in hardware, networks, or systems.
- Respond to alerts and incidents to minimize disruptions and maintain system uptime.
- Ensure high availability of systems in 24×7 production environments.
Performance Validation & Optimization¶
- Run performance benchmarks using tools like HPL, LINPACK, and HPCC.
- Validate performance after system upgrades, changes, or new deployments.
- Work on fine-tuning system performance to meet operational goals.
Application & User Support¶
- Install, configure, and validate both commercial and open-source applications in HPC environments.
- Support running applications that use MPI for multi-node execution.
- Manage environment modules and application runtime settings to ensure smooth operation.
- Provide user onboarding, access management, and technical support to researchers and engineers.
Automation & Operational Efficiency¶
- Write shell scripts to automate common administrative tasks.
- Automate processes like OS provisioning, user management, and software deployment.
- Streamline operations to improve consistency and reduce manual effort.
Hardware & Infrastructure Support¶
- Assist with server installation, integration, and validation in datacenter environments.
- Perform BIOS, firmware, and driver upgrades on Dell PowerEdge platforms.
- Troubleshoot hardware issues and assist with diagnostics and repairs.
- Help coordinate infrastructure changes during upgrades and system expansions.
Networking Support¶
- Manage high-speed InfiniBand (400 Gb/s) networking for HPC clusters.
- Monitor network health, fix connectivity issues, and optimize performance.
- Ensure network readiness for running large-scale parallel workloads.
Incident Management & Documentation¶
- Provide L2-level incident support in high-availability environments.
- Analyze and troubleshoot system issues, then take corrective actions.
- Document incidents, solutions, and operational procedures for future reference.
- Support activities related to system changes and updates.
Datacenter & Project Support¶
- Contribute to projects aimed at datacenter modernization and infrastructure deployment.
- Support tasks like server installation, OS provisioning, and network setup.
- Work closely with cross-functional teams during project execution and handover.