Skip to content

Roles & Responsibilities

This section outlines my key roles and responsibilities in supporting large-scale Linux and high-performance computing (HPC) environments in production settings.

HPC Operations & Administration

  • Manage and support multi-petaflop HPC environments for research workloads.
  • Perform tasks like setting up and managing compute nodes, from installation to decommissioning.
  • Support workload scheduling using PBS Pro to ensure tasks run smoothly in shared environments.
  • Ensure stable job execution across both CPU and GPU-based clusters.

System Monitoring & Reliability

  • Monitor system health and availability using tools like NVIDIA Bright Cluster Manager, Dell OpenManage Enterprise, and InfiniBand UFM.
  • Run proactive health checks to identify potential issues in hardware, networks, or systems.
  • Respond to alerts and incidents to minimize disruptions and maintain system uptime.
  • Ensure high availability of systems in 24×7 production environments.

Performance Validation & Optimization

  • Run performance benchmarks using tools like HPL, LINPACK, and HPCC.
  • Validate performance after system upgrades, changes, or new deployments.
  • Work on fine-tuning system performance to meet operational goals.

Application & User Support

  • Install, configure, and validate both commercial and open-source applications in HPC environments.
  • Support running applications that use MPI for multi-node execution.
  • Manage environment modules and application runtime settings to ensure smooth operation.
  • Provide user onboarding, access management, and technical support to researchers and engineers.

Automation & Operational Efficiency

  • Write shell scripts to automate common administrative tasks.
  • Automate processes like OS provisioning, user management, and software deployment.
  • Streamline operations to improve consistency and reduce manual effort.

Hardware & Infrastructure Support

  • Assist with server installation, integration, and validation in datacenter environments.
  • Perform BIOS, firmware, and driver upgrades on Dell PowerEdge platforms.
  • Troubleshoot hardware issues and assist with diagnostics and repairs.
  • Help coordinate infrastructure changes during upgrades and system expansions.

Networking Support

  • Manage high-speed InfiniBand (400 Gb/s) networking for HPC clusters.
  • Monitor network health, fix connectivity issues, and optimize performance.
  • Ensure network readiness for running large-scale parallel workloads.

Incident Management & Documentation

  • Provide L2-level incident support in high-availability environments.
  • Analyze and troubleshoot system issues, then take corrective actions.
  • Document incidents, solutions, and operational procedures for future reference.
  • Support activities related to system changes and updates.

Datacenter & Project Support

  • Contribute to projects aimed at datacenter modernization and infrastructure deployment.
  • Support tasks like server installation, OS provisioning, and network setup.
  • Work closely with cross-functional teams during project execution and handover.