Hi I am
Ayyappaswamy Eethakota
I work on high-performance computing systems — building and optimizing the infrastructure that powers demanding, large-scale workloads.
I am Technology Associate in High-Performance Computing (HPC) with hands-on experience managing large-scale, high-performance computing environments for government and enterprise research workloads. My work centers on keeping production-grade HPC infrastructure reliable, performant, and continuously available.
I began my HPC career in January 2024 and have since built expertise across Linux system administration, cluster deployment and operations, server provisioning, job scheduling, high-speed networking, parallel file systems, GPU infrastructure, and performance troubleshooting. I currently serve as a Technology Associate – HPC at my second organization, where I design, build, and manage HPC environments that scale.
A core part of my experience has been contributing to the deployment and ongoing operations of India's 11-petaflop supercomputing infrastructure — an environment spanning over 1,300 compute nodes and NVIDIA H100 GPU accelerator systems. Day to day, this means provisioning clusters, managing operating systems, scheduling workloads, monitoring system health, and troubleshooting across both CPU and GPU nodes to minimize downtime. I also manage the Lustre file system, ensuring efficient, high-throughput data storage and access across the environment.
Beyond infrastructure operations, I'm continually deepening my skills in automation, storage architecture, networking, and data-center operations — with the goal of not just maintaining HPC systems, but making them faster, more resilient, and easier to operate at scale.
Managing the provisioning, configuration, and maintenance of operating systems across compute nodes, ensuring proper setup, updates, and lifecycle management throughout each node's usage. This includes integrating hardware and software so the entire system runs efficiently.
Responsible for maintaining and troubleshooting hardware components such as CPUs, memory modules, network interface cards (NICs), and storage devices. I diagnose hardware failures, perform component replacements, and ensure optimal hardware functionality across the entire cluster.
Configuring and optimizing PBS Pro for efficient workload scheduling, ensuring smooth execution of tasks across both CPU and GPU-based clusters. I fine-tune resource allocation to maximize system performance and minimize job queuing delays.
Maintaining the Lustre file system to ensure high-performance, scalable storage across multiple compute nodes. I troubleshoot data access issues, optimize file system performance, and ensure data integrity and availability for research workloads.
Monitoring system performance in real time, validating hardware and software health, and addressing potential issues proactively. I use tools like Dell OpenManage Enterprise (OME) and Bright Cluster Manager for hardware and system health validation, ensuring the cluster remains fully operational.
Managing user access and permissions within the HPC environment while ensuring security protocols are followed. I troubleshoot system-level issues related to hardware, software, and network connectivity to minimize disruptions to research workflows.
Installing and configuring both commercial and open-source applications on the cluster, ensuring all required dependencies are met. I manage essential libraries such as MPI, CUDA, and OpenCL and resolve compatibility issues so applications run smoothly and efficiently.
Utilizing platforms like Bright Cluster Manager, Dell OpenManage Enterprise (OME), and InfiniBand UFM to monitor system performance, track hardware health, and resolve issues early. This supports a high system availability target of 99.99% uptime and optimized performance for distributed workloads.
Supporting InfiniBand networks by monitoring link health, troubleshooting connectivity issues, and assisting with network optimization. I help ensure network performance remains strong, which is essential for high-throughput data transfer in HPC environments.
Automation plays a crucial role in enhancing efficiency and reducing manual effort across server management. I use shell scripting to automate several key operations, including:
Automating server configurations to ensure uniformity across all systems, including managing hardware drivers and system settings, and verifying that every server is configured consistently with required parameters.
Automating server monitoring tasks, including collecting server data such as CPU and memory usage and checking the status of critical services, so I can quickly identify potential issues such as resource exhaustion or service failures and take action to prevent disruptions.
Writing scripts to identify hardware or driver mismatches across servers, ensuring correct hardware driver configurations and preventing inconsistencies that could affect performance or cause errors.
I provide 24×7 L2 operational support in high-availability environments, focusing on:
My primary goal is to ensure a stable infrastructure, predictable performance, and a smooth experience for researchers and engineers working on compute-intensive workloads — maintaining uptime and reliability across all systems.
My approach centers on three key areas: stability, performance, and scalability. I work to build and maintain Linux and HPC environments that are:
Strong, secure, and consistently high-performing.
Closely watched for early detection of issues.
Built to handle scientific and engineering workloads at the scale complex computing requires.
My technical skill set covers Linux system administration, high-performance computing (HPC), software tools, networking, hardware management, and data management within enterprise environments.
This section outlines my key roles and responsibilities in supporting large-scale Linux and high-performance computing (HPC) environments in production settings.
This section lists the commercial and open-source software, libraries, compilers, and tools I have worked with in high-performance computing environments.
Installation and validation on multi-node HPC clusters, MPI and scheduler integration, environment module creation, batch and interactive job support, and runtime troubleshooting.
Source-based installation, configuration, MPI integration, multi-node execution testing, environment setup, and troubleshooting build, runtime, and scaling issues.
Application compilation, build validation, library and compiler compatibility checks, runtime library configuration, and support for CPU and GPU-enabled workloads.
I was part of the core team responsible for deploying and maintaining a multi-petaflops high-performance computing system used for national research and government tasks.
I contributed to a major datacenter modernization project to support enterprise and HPC-ready infrastructure.