Hi I am

Ayyappaswamy Eethakota

Data Centre
& HPC Engineer

I work on high-performance computing systems — building and optimizing the infrastructure that powers demanding, large-scale workloads.

2+
Years Experience
1300+
Node Cluster
13 PB
Storage
Ayyappaswamy Eethakota

Professional Summary

I am Technology Associate in High-Performance Computing (HPC) with hands-on experience managing large-scale, high-performance computing environments for government and enterprise research workloads. My work centers on keeping production-grade HPC infrastructure reliable, performant, and continuously available.

I began my HPC career in January 2024 and have since built expertise across Linux system administration, cluster deployment and operations, server provisioning, job scheduling, high-speed networking, parallel file systems, GPU infrastructure, and performance troubleshooting. I currently serve as a Technology Associate – HPC at my second organization, where I design, build, and manage HPC environments that scale.

A core part of my experience has been contributing to the deployment and ongoing operations of India's 11-petaflop supercomputing infrastructure — an environment spanning over 1,300 compute nodes and NVIDIA H100 GPU accelerator systems. Day to day, this means provisioning clusters, managing operating systems, scheduling workloads, monitoring system health, and troubleshooting across both CPU and GPU nodes to minimize downtime. I also manage the Lustre file system, ensuring efficient, high-throughput data storage and access across the environment.

Beyond infrastructure operations, I'm continually deepening my skills in automation, storage architecture, networking, and data-center operations — with the goal of not just maintaining HPC systems, but making them faster, more resilient, and easier to operate at scale.

Core Responsibilities

OS Provisioning and Node Lifecycle Management

Managing the provisioning, configuration, and maintenance of operating systems across compute nodes, ensuring proper setup, updates, and lifecycle management throughout each node's usage. This includes integrating hardware and software so the entire system runs efficiently.

Hardware Management and Troubleshooting

Responsible for maintaining and troubleshooting hardware components such as CPUs, memory modules, network interface cards (NICs), and storage devices. I diagnose hardware failures, perform component replacements, and ensure optimal hardware functionality across the entire cluster.

Scheduler Integration and Workload Execution Using PBS Pro

Configuring and optimizing PBS Pro for efficient workload scheduling, ensuring smooth execution of tasks across both CPU and GPU-based clusters. I fine-tune resource allocation to maximize system performance and minimize job queuing delays.

Lustre File System Management and Troubleshooting

Maintaining the Lustre file system to ensure high-performance, scalable storage across multiple compute nodes. I troubleshoot data access issues, optimize file system performance, and ensure data integrity and availability for research workloads.

Continuous Monitoring and Health Validation

Monitoring system performance in real time, validating hardware and software health, and addressing potential issues proactively. I use tools like Dell OpenManage Enterprise (OME) and Bright Cluster Manager for hardware and system health validation, ensuring the cluster remains fully operational.

User Access Control and System-Level Troubleshooting

Managing user access and permissions within the HPC environment while ensuring security protocols are followed. I troubleshoot system-level issues related to hardware, software, and network connectivity to minimize disruptions to research workflows.

Application Installation and Library Management

Installing and configuring both commercial and open-source applications on the cluster, ensuring all required dependencies are met. I manage essential libraries such as MPI, CUDA, and OpenCL and resolve compatibility issues so applications run smoothly and efficiently.

System Monitoring Tools

Utilizing platforms like Bright Cluster Manager, Dell OpenManage Enterprise (OME), and InfiniBand UFM to monitor system performance, track hardware health, and resolve issues early. This supports a high system availability target of 99.99% uptime and optimized performance for distributed workloads.

High-Speed InfiniBand (400 Gb/s) Network Support

Supporting InfiniBand networks by monitoring link health, troubleshooting connectivity issues, and assisting with network optimization. I help ensure network performance remains strong, which is essential for high-throughput data transfer in HPC environments.

Automation & Operations

Automation plays a crucial role in enhancing efficiency and reducing manual effort across server management. I use shell scripting to automate several key operations, including:

Server Configuration and Actions

Automating server configurations to ensure uniformity across all systems, including managing hardware drivers and system settings, and verifying that every server is configured consistently with required parameters.

Server Monitoring and Data Collection

Automating server monitoring tasks, including collecting server data such as CPU and memory usage and checking the status of critical services, so I can quickly identify potential issues such as resource exhaustion or service failures and take action to prevent disruptions.

Identifying Configuration Issues

Writing scripts to identify hardware or driver mismatches across servers, ensuring correct hardware driver configurations and preventing inconsistencies that could affect performance or cause errors.

Operations & Support

I provide 24×7 L2 operational support in high-availability environments, focusing on:

  • Incident diagnosis and root-cause analysis
  • Corrective actions and continuous system optimization
  • Comprehensive documentation of troubleshooting processes and solutions

My primary goal is to ensure a stable infrastructure, predictable performance, and a smooth experience for researchers and engineers working on compute-intensive workloads — maintaining uptime and reliability across all systems.

Focus & Approach

My approach centers on three key areas: stability, performance, and scalability. I work to build and maintain Linux and HPC environments that are:

Reliable

Strong, secure, and consistently high-performing.

Well-Monitored

Closely watched for early detection of issues.

Ready for Demanding Tasks

Built to handle scientific and engineering workloads at the scale complex computing requires.

Technical Skills

My technical skill set covers Linux system administration, high-performance computing (HPC), software tools, networking, hardware management, and data management within enterprise environments.

Operating Systems

  • Red Hat Enterprise Linux (RHEL) 9.x
  • CentOS
  • Linux system administration in multi-node environments
  • OS installation and configuration

HPC Cluster Management & Scheduling

  • NVIDIA Bright Cluster Manager
  • PBS Pro
  • Slurm

Activities Include

  • Cluster provisioning
  • Node lifecycle management
  • Scheduler integration
  • Job monitoring
  • Production troubleshooting

Networking & Interconnects

  • InfiniBand (QM9700 / QM9790 – 400 Gb/s)
  • Dell S3100 Ethernet switches (1G / 10G)
  • InfiniBand fabric monitoring and issue resolution
  • Network validation for HPC performance workloads

Monitoring & Management Tools

  • NVIDIA Bright Cluster Manager
  • Dell OpenManage Enterprise (OME)
  • InfiniBand UFM

Used For

  • System health monitoring
  • Proactive fault detection
  • Hardware alerts
  • Operational reliability

Benchmarking & Performance Validation

  • HPL
  • HPCG
  • HPCC
  • LAMMPS

Experience Includes

  • Performance validation
  • Baseline comparison
  • Throughput verification
  • Deployment and production testing

Package & Configuration Management

  • YUM
  • DNF
  • RPM
  • Repository configuration
  • Dependency resolution in HPC environments

Storage & File Systems

  • Parallel File Systems: Lustre, BeeGFS — deployment, tuning, and troubleshooting
  • Volume Management: LVM, NFS

GPU & Accelerated Computing

  • GPU Infrastructure: NVIDIA H100 accelerator systems, CUDA, driver and library management

Hardware & Datacenter Operations

  • Dell PowerEdge servers: R6625, R7625, XE8640, T550
  • BIOS, firmware, and driver upgrades
  • Hardware diagnostics and fault handling
  • Server integration and validation

Operations & Support

  • 24×7 L2 production support
  • Incident diagnosis and root-cause analysis
  • Change implementation and validation
  • Operational documentation and handover support

Professional Experience

Technology Associate in High-Performance Computing (HPC)

January 2024 – Present

My HPC career began in January 2024, during which I have worked across various organizations on the administration, operation, and reliability of large-scale Linux and High-Performance Computing (HPC) environments supporting mission-critical overnment and enterprise research workloads.

My role involves day-to-day production support, cluster operations, performance validation, and infrastructure maintenance across CPU and GPU-based HPC systems..

Key Responsibilities & Contributions

  • Administered and supported large-scale HPC clusters comprising over 1,300+ compute nodes and GPU accelerator nodes with NVIDIA H100 GPUs, sustaining production workloads at multi-petaflop scale.
  • Managed HPC environments using NVIDIA Bright Cluster Manager, including cluster provisioning, node lifecycle management, configuration updates, and system validation.
  • Supported workload scheduling and job execution using PBS Pro, ensuring stable and predictable job execution in shared environments.
  • Performed continuous system monitoring and health checks using Bright Cluster Manager, Dell OpenManage Enterprise, and InfiniBand UFM, contributing to 99.99% system availability.
  • Executed OS provisioning, patching, and configuration management across compute and management nodes.
  • Automated routine administrative tasks using shell scripting, improving operational efficiency and reducing manual intervention.
  • Installed, configured, and validated commercial and open-source HPC applications across multi-node environments, including MPI integration and environment module configuration.
  • Conducted performance benchmarking and validation using HPL, LINPACK, and HPCC.
  • Supported 400 Gb/s InfiniBand networking, including fabric monitoring, connectivity troubleshooting, and performance tuning.
  • Performed hardware-level operations including BIOS, firmware, and driver upgrades on Dell PowerEdge servers.
  • Provided 24×7 L2 operational support, handling incident diagnosis, root-cause analysis, corrective actions, and operational documentation.
  • Contributed to datacenter modernization activities, supporting server integration, network validation, and operating system provisioning.

This role requires close coordination with infrastructure, networking, and application teams to ensure stable, secure, and high-performing HPC environments for production use.

Roles & Responsibilities

This section outlines my key roles and responsibilities in supporting large-scale Linux and high-performance computing (HPC) environments in production settings.

HPC Operations & Administration

  • Manage and support multi-petaflop HPC environments for research workloads.
  • Perform tasks like setting up and managing compute nodes, from installation to decommissioning.
  • Support workload scheduling using PBS Pro.
  • Ensure stable job execution across CPU and GPU-based clusters.

System Monitoring & Reliability

  • Monitor system health and availability using NVIDIA Bright Cluster Manager, Dell OpenManage Enterprise, and InfiniBand UFM.
  • Run proactive health checks to identify potential issues.
  • Respond to alerts and incidents to minimize disruptions.
  • Ensure high availability of systems in 24×7 production environments.

Performance Validation & Optimization

  • Run performance benchmarks using HPL, LINPACK, and HPCC.
  • Validate performance after system upgrades, changes, or new deployments.
  • Work on fine-tuning system performance to meet operational goals.

Application & User Support

  • Install, configure, and validate commercial and open-source applications.
  • Support applications using MPI for multi-node execution.
  • Manage environment modules and application runtime settings.
  • Provide user onboarding, access management, and technical support.

Automation & Operational Efficiency

  • Write shell scripts to automate common administrative tasks.
  • Automate OS provisioning, user management, and software deployment.
  • Streamline operations to improve consistency and reduce manual effort.

Hardware & Infrastructure Support

  • Assist with server installation, integration, and validation.
  • Perform BIOS, firmware, and driver upgrades on Dell PowerEdge platforms.
  • Troubleshoot hardware issues and assist with diagnostics and repairs.
  • Help coordinate infrastructure changes during upgrades and system expansions.

Networking Support

  • Manage high-speed InfiniBand (400 Gb/s) networking for HPC clusters.
  • Monitor network health, fix connectivity issues, and optimize performance.
  • Ensure network readiness for large-scale parallel workloads.

Incident Management & Documentation

  • Provide L2-level incident support in high-availability environments.
  • Analyze and troubleshoot system issues, then take corrective actions.
  • Document incidents, solutions, and operational procedures.
  • Support activities related to system changes and updates.

Datacenter & Project Support

  • Contribute to datacenter modernization and infrastructure deployment.
  • Support server installation, OS provisioning, and network setup.
  • Work closely with cross-functional teams during project execution and handover.

Applications & Libraries

This section lists the commercial and open-source software, libraries, compilers, and tools I have worked with in high-performance computing environments.

Commercial Applications

  • ANSYS Fluent
  • ANSYS CFX
  • STAR-CCM+
  • HEEDS
  • Altair FEKO

Installation and validation on multi-node HPC clusters, MPI and scheduler integration, environment module creation, batch and interactive job support, and runtime troubleshooting.

Open-Source Applications

  • OpenFOAM
  • SU2
  • ParaView
  • WRF

Source-based installation, configuration, MPI integration, multi-node execution testing, environment setup, and troubleshooting build, runtime, and scaling issues.

Compilers, Libraries & Toolchains

Compilers

  • AOCC — AMD Optimizing C/C++ Compiler
  • GCC
  • Intel OneAPI Toolkit

Libraries & MPI

  • AOCL — AMD Optimized CPU Libraries
  • OpenMPI

Application compilation, build validation, library and compiler compatibility checks, runtime library configuration, and support for CPU and GPU-enabled workloads.

Operational Scope

  • Installation and upgrade validation
  • Dependency and compatibility checks
  • Scheduler and resource manager integration
  • Environment module management
  • User support and issue resolution
  • Coordination with infrastructure and application teams

Projects

India’s Multi-Petaflops HPC Cluster Deployment

I was part of the core team responsible for deploying and maintaining a multi-petaflops high-performance computing system used for national research and government tasks.

Scope & Contributions

  • Supported an HPC environment delivering about 11 Petaflops of compute power.
  • Managed clusters with over 1,300+ compute nodes and NVIDIA H100 GPUs.
  • Managed 13 Petabytes of Lustre storage, including monitoring and troubleshooting.
  • Supported 30 Petabytes of tape library storage for data archiving and retrieval.
  • Set up and configured the cluster using NVIDIA Bright Cluster Manager.
  • Managed job scheduling using PBS Pro.
  • Validated system performance using HPL, LINPACK, and HPCC.
  • Monitored and troubleshot 400 Gb/s InfiniBand fabric during cluster operations.

Datacenter Infrastructure Modernization

I contributed to a major datacenter modernization project to support enterprise and HPC-ready infrastructure.

  • Supported server installation and integration.
  • Handled operating system installation and validation.
  • Assisted with network connectivity tests and readiness checks.
  • Supported hardware configuration and firmware validation.
  • Worked with infrastructure and networking teams during deployment.

Ongoing Operational Projects

  • Cluster expansion and adding new nodes
  • Hardware refreshes and firmware upgrades
  • Performance validation after system updates
  • Automation of provisioning and administration processes
  • Incident resolution and continuous system improvement