Skip to content

Professional Experience

Technology Associate in High-Performance Computing (HPC)

January 2024 – Present

My HPC career began in January 2024, during which I have worked across various organizations on the administration, operation, and reliability of large-scale Linux and High-Performance Computing (HPC) environments supporting mission-critical government and enterprise research workloads.

My role involves day-to-day production support, cluster operations, performance validation, and infrastructure maintenance across CPU and GPU-based HPC systems.

Key Responsibilities & Contributions

  • Administered and supported large-scale HPC clusters comprising over 1,300+ compute nodes and GPU accelerator nodes with NVIDIA H100 GPUs, sustaining production workloads at multi-petaflop scale.
  • Managed HPC environments using NVIDIA Bright Cluster Manager, including cluster provisioning, node lifecycle management, configuration updates, and system validation.
  • Supported workload scheduling and job execution using PBS Pro, ensuring stable and predictable job execution in shared environments.
  • Performed continuous system monitoring and health checks using Bright Cluster Manager, Dell OpenManage Enterprise, and InfiniBand UFM, contributing to 99.99% system availability.
  • Executed OS provisioning, patching, and configuration management across compute and management nodes in production environments.
  • Automated routine administrative tasks using shell scripting, improving operational efficiency and reducing manual intervention.
  • Installed, configured, and validated commercial and open-source HPC applications across multi-node environments, including MPI integration and environment module configuration.
  • Conducted performance benchmarking and validation using HPL, LINPACK, and HPCC, ensuring systems met expected throughput and floating-point performance targets.
  • Supported 400 Gb/s InfiniBand networking, including fabric monitoring, troubleshooting connectivity issues, and assisting with performance tuning.
  • Performed hardware-level operations, including BIOS, firmware, and driver upgrades on Dell PowerEdge servers, ensuring compliance with operational and security standards.
  • Provided 24×7 L2 operational support, handling incident diagnosis, root-cause analysis, corrective actions, and operational documentation.
  • Contributed to datacenter modernization activities, supporting server integration, network validation, and operating system provisioning across multi-node environments.

This role requires close coordination with infrastructure, networking, and application teams to ensure stable, secure, and high-performing HPC environments for production use.