Projects¶
India’s Multi-Petaflops HPC Cluster Deployment¶
I was part of the core team responsible for deploying and maintaining a multi-petaflops high-performance computing system used for national research and government tasks.
Scope & Contributions:
- Supported an HPC environment delivering about 11 Petaflops of compute power.
- Managed clusters with over 1,300+ compute nodes and NVIDIA H100 GPUs.
- Managed 13 Petabytes of Lustre storage, including monitoring and troubleshooting for optimal performance.
- Supported 30 Petabytes of tape library storage to ensure secure and efficient data archiving and retrieval.
- Set up and configure the cluster using NVIDIA Bright Cluster Manager.
- Managed job scheduling and task execution using PBS Pro in shared environments.
- Validated system performance using HPL, LINPACK, and HPCC to ensure stable compute throughput.
- Participated in checks before production launch, confirming that compute, network, and monitoring systems were fully operational.
- Monitored and troubleshot 400 Gb/s InfiniBand fabric during cluster operations.
- This project required high standards of performance, availability, and reliability for national research workloads.
Datacenter Infrastructure Modernization¶
I contributed to a major datacenter modernization project to support enterprise and HPC-ready infrastructure.
Scope & Contributions:
- Supported server installation and integration in the datacenter environment.
- Handled operating system installation and validation across several nodes.
- Assisted with network connectivity tests and system readiness checks.
- Supported hardware configuration, firmware validation, and initial system setup.
- Worked closely with infrastructure and networking teams to ensure smooth deployment and transition to production.
- This project aimed to improve infrastructure reliability, scalability, and readiness for HPC and enterprise workloads.
Ongoing Operational Projects¶
Along with large deployments, I am involved in several ongoing operational projects, including:
- Cluster expansion and adding new nodes
- Hardware refreshes and firmware upgrades
- Performance validation after system updates or changes
- Automating processes for provisioning and administration
- Addressing incidents and improving systems based on feedback