Skip to content

Projects


India’s Multi-Petaflops HPC Cluster Deployment

I was part of the core team responsible for deploying and maintaining a multi-petaflops high-performance computing system used for national research and government tasks.

Scope & Contributions:

  • Supported an HPC environment delivering about 11 Petaflops of compute power.
  • Managed clusters with over 1,300+ compute nodes and NVIDIA H100 GPUs.
  • Managed 13 Petabytes of Lustre storage, including monitoring and troubleshooting for optimal performance.
  • Supported 30 Petabytes of tape library storage to ensure secure and efficient data archiving and retrieval.
  • Set up and configure the cluster using NVIDIA Bright Cluster Manager.
  • Managed job scheduling and task execution using PBS Pro in shared environments.
  • Validated system performance using HPL, LINPACK, and HPCC to ensure stable compute throughput.
  • Participated in checks before production launch, confirming that compute, network, and monitoring systems were fully operational.
  • Monitored and troubleshot 400 Gb/s InfiniBand fabric during cluster operations.
  • This project required high standards of performance, availability, and reliability for national research workloads.

Datacenter Infrastructure Modernization

I contributed to a major datacenter modernization project to support enterprise and HPC-ready infrastructure.

Scope & Contributions:

  • Supported server installation and integration in the datacenter environment.
  • Handled operating system installation and validation across several nodes.
  • Assisted with network connectivity tests and system readiness checks.
  • Supported hardware configuration, firmware validation, and initial system setup.
  • Worked closely with infrastructure and networking teams to ensure smooth deployment and transition to production.
  • This project aimed to improve infrastructure reliability, scalability, and readiness for HPC and enterprise workloads.

Ongoing Operational Projects

Along with large deployments, I am involved in several ongoing operational projects, including:

  • Cluster expansion and adding new nodes
  • Hardware refreshes and firmware upgrades
  • Performance validation after system updates or changes
  • Automating processes for provisioning and administration
  • Addressing incidents and improving systems based on feedback