Professional Experience¶
Technology Associate in High-Performance Computing (HPC)¶
January 2024 – Present
My HPC career began in January 2024, during which I have worked across various organizations on the administration, operation, and reliability of large-scale Linux and High-Performance Computing (HPC) environments supporting mission-critical government and enterprise research workloads.
My role involves day-to-day production support, cluster operations, performance validation, and infrastructure maintenance across CPU and GPU-based HPC systems.
Key Responsibilities & Contributions¶
- Administered and supported large-scale HPC clusters comprising over 1,300+ compute nodes and GPU accelerator nodes with NVIDIA H100 GPUs, sustaining production workloads at multi-petaflop scale.
- Managed HPC environments using NVIDIA Bright Cluster Manager, including cluster provisioning, node lifecycle management, configuration updates, and system validation.
- Supported workload scheduling and job execution using PBS Pro, ensuring stable and predictable job execution in shared environments.
- Performed continuous system monitoring and health checks using Bright Cluster Manager, Dell OpenManage Enterprise, and InfiniBand UFM, contributing to 99.99% system availability.
- Executed OS provisioning, patching, and configuration management across compute and management nodes in production environments.
- Automated routine administrative tasks using shell scripting, improving operational efficiency and reducing manual intervention.
- Installed, configured, and validated commercial and open-source HPC applications across multi-node environments, including MPI integration and environment module configuration.
- Conducted performance benchmarking and validation using HPL, LINPACK, and HPCC, ensuring systems met expected throughput and floating-point performance targets.
- Supported 400 Gb/s InfiniBand networking, including fabric monitoring, troubleshooting connectivity issues, and assisting with performance tuning.
- Performed hardware-level operations, including BIOS, firmware, and driver upgrades on Dell PowerEdge servers, ensuring compliance with operational and security standards.
- Provided 24×7 L2 operational support, handling incident diagnosis, root-cause analysis, corrective actions, and operational documentation.
- Contributed to datacenter modernization activities, supporting server integration, network validation, and operating system provisioning across multi-node environments.
This role requires close coordination with infrastructure, networking, and application teams to ensure stable, secure, and high-performing HPC environments for production use.