Skip to content

Professional Summary

I am a Linux Systems Administrator and HPC Engineer with hands-on experience managing large-scale, high-performance computing environments for government and enterprise research workloads. My work centers on keeping production-grade HPC infrastructure reliable, performant, and continuously available.

I began my HPC career in January 2024 and have since built expertise across Linux system administration, cluster deployment and operations, server provisioning, job scheduling, high-speed networking, parallel file systems, GPU infrastructure, and performance troubleshooting. I currently serve as a Technology Associate – HPC at my second organization, where I design, build, and manage HPC environments that scale.

A core part of my experience has been contributing to the deployment and ongoing operations of India's 11-petaflop supercomputing infrastructure — an environment spanning over 1,300 compute nodes and NVIDIA H100 GPU accelerator systems. Day to day, this means provisioning clusters, managing operating systems, scheduling workloads, monitoring system health, and troubleshooting across both CPU and GPU nodes to minimize downtime. I also manage the Lustre file system, ensuring efficient, high-throughput data storage and access across the environment.

Beyond infrastructure operations, I'm continually deepening my skills in automation, storage architecture, networking, and data-center operations — with the goal of not just maintaining HPC systems, but making them faster, more resilient, and easier to operate at scale.


Core Responsibilities

OS Provisioning and Node Lifecycle Management Managing the provisioning, configuration, and maintenance of operating systems across compute nodes, ensuring proper setup, updates, and lifecycle management throughout each node's usage. This includes integrating hardware and software so the entire system runs efficiently.

Hardware Management and Troubleshooting Responsible for maintaining and troubleshooting hardware components such as CPUs, memory modules, network interface cards (NICs), and storage devices. I diagnose hardware failures, perform component replacements, and ensure optimal hardware functionality across the entire cluster.

Scheduler Integration and Workload Execution Using PBS Pro Configuring and optimizing PBS Pro for efficient workload scheduling, ensuring smooth execution of tasks across both CPU and GPU-based clusters. I fine-tune resource allocation to maximize system performance and minimize job queuing delays.

Lustre File System Management and Troubleshooting Maintaining the Lustre file system to ensure high-performance, scalable storage across multiple compute nodes. I troubleshoot data access issues, optimize file system performance, and ensure data integrity and availability for research workloads.

Continuous Monitoring and Health Validation Monitoring system performance in real time, validating hardware and software health, and addressing potential issues proactively. I use tools like Dell OpenManage Enterprise (OME) and Bright Cluster Manager for hardware and system health validation, ensuring the cluster remains fully operational.

User Access Control and System-Level Troubleshooting Managing user access and permissions within the HPC environment while ensuring security protocols are followed. I troubleshoot system-level issues related to hardware, software, and network connectivity to minimize disruptions to research workflows.

Application Installation and Library Management Installing and configuring both commercial and open-source applications on the cluster, ensuring all required dependencies are met. I manage essential libraries (e.g., MPI, CUDA, OpenCL) and resolve compatibility issues so applications run smoothly and efficiently.

System Monitoring Tools Utilizing platforms like Bright Cluster Manager, Dell OpenManage Enterprise (OME), and InfiniBand UFM to monitor system performance, track hardware health, and resolve issues early. This supports a high system availability target of 99.99% uptime and optimized performance for distributed workloads.

High-Speed InfiniBand (400 Gb/s) Network Support Supporting InfiniBand networks by monitoring link health, troubleshooting connectivity issues, and assisting with network optimization. I help ensure network performance remains strong, which is essential for high-throughput data transfer in HPC environments.


Automation & Operations

Automation plays a crucial role in enhancing efficiency and reducing manual effort across server management. I use shell scripting to automate several key operations, including:

  • Server Configuration and Actions — Automating server configurations to ensure uniformity across all systems, including managing hardware drivers and system settings, and verifying that every server is configured consistently with required parameters.

  • Server Monitoring and Data Collection — Automating server monitoring tasks, including collecting server data (CPU, memory) and checking the status of critical services, so I can quickly identify potential issues such as resource exhaustion or service failures and take action to prevent disruptions.

  • Identifying Configuration Issues — Writing scripts to identify hardware or driver mismatches across servers, ensuring correct hardware driver configurations and preventing inconsistencies that could affect performance or cause errors.


Operations & Support

I provide 24×7 L2 operational support in high-availability environments, focusing on:

  • Incident diagnosis and root-cause analysis
  • Corrective actions and continuous system optimization
  • Comprehensive documentation of troubleshooting processes and solutions

My primary goal is to ensure a stable infrastructure, predictable performance, and a smooth experience for researchers and engineers working on compute-intensive workloads — maintaining uptime and reliability across all systems.


Focus & Approach

My approach centers on three key areas: stability, performance, and scalability. I work to build and maintain Linux and HPC environments that are:

  • Reliable — strong, secure, and consistently high-performing.
  • Well-monitored — closely watched for early detection of issues.
  • Ready for demanding tasks — built to handle scientific and engineering workloads at the scale complex computing requires.