====== HPC Newsletter - 2026/05 - 06 ====== Welcome to the May/June 2026 HPC newsletter. This is an unusual update combining both May and June changes due to the extended service outage of Comet which began at the end of May. ===== Join us for the Research Computing Community Kick-off Workshop ===== **Monday 29th June**, Henry Daysh Building Room 1.04, 14:30-15:30 Are you writing or adapting code for your research, using data-intensive methods, or high-performance computing (HPC)? We invite you to join an initial in-person workshop to help build a new community of practice. This informal 1 hour session will be an opportunity to meet colleagues, share experiences and challenges, explore common interests, and help shape future activities and support around research coding, data, and HPC. Whether you're just getting started with coding and data for your research, or have experience to share, we'd love to hear your perspective and have you involved from the start. **Register or express interest here: [[https://forms.office.com/e/VTczQGVPs6|HPC and Code Community Kick off Meeting – Fill in form]]** ==== Upcoming Training Sessions ==== * Book [[https://pretix.eu/ncl/2026-06-23-NCL/| Version Control with Git]] - **23/06/2026** - PGRLL6.19 * Book [[https://pretix.eu/ncl on 02/07/2026 |HPC Induction]] - **02/07/2026** - Online/Teams * Book [[https://pretix.eu/ncl/2026-07-09-NCL|Intro to HPC workshop]] - **09/07/2026** - HDB1.14.PC * Book [[https://pretix.eu/ncl/2026-07-21-NCL/|Programming with R workshop]] - **21/07/2026** - PGRLL6.19 * Book [[https://pretix.eu/ncl/2026-07-30-NCL/|Intro to HPC workshop]] - **30/07/2026** - PGRLL6.19 ===== HPC Summary for May-June 2026 ===== * Registered projects: **565** * Active projects: **202** * Total users: **1045** * Total driving tests: **1130** tests taken, **456** passes * CPU time: hours of compute May - June : 2526336 ===== Software Changes ===== **New software** * Amrfinderplus, Antismash, Convert3D, Eggnog, NCBI Blast+, Meme, Miniphy, Phyalign, Diamond and more added to the Bioapps container - some of these are already installed as (older) modules but we intend to //not// update those modules in the future, instead using the Bioapps container for most updated going forward * Central installation of many [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=advanced:software:bioapps#application_databases_data_files|popular Bioinformatics databases]] to support the new software - 4TB of data in a shared location (/nobackups/shared/data) accessible by all users * Release of our own [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=started:simple_slurm_tools|Simple Slurm Tools]], which you can use to monitor queues, query your project limits/resources, etc * [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/dokuwiki/doku.php?id=advanced:software:gipsyx|GipsyX]] added - access is restricted to specific project groups **Changed software** * [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=advanced:software:ansys|ANSYS 2025 R2]] - Now migrating to a container environment, away from the old module based versions * [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=advanced:matlab|Matlab 2026a]] - Default version now updated to 2026 * [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=advanced:software:bioapps|Bioapps container]] image updated to 2026.06 ===== Website & Documentation ===== * All installed [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=advanced:software_list|software]] now has a guide page in the wiki * Start of our new [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=advanced:slurm|Advanced Slurm Job Types]] documentation which will outline the two main approaches to parallelism on HPC (task arrays and MPI). These build on the simple examples we introduce at the end of the 'Introduction to HPC' entry workshops for new HPC users. * [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=started:job_parallel|How to build task arrays]] - complete guide available and ready to use * [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/dokuwiki/doku.php?id=advanced:slurm_mpi_example|How to build simple MPI solutions]] - guide in-progress ===== System Changes ===== As you will all be aware, we had an unplanned extended loss of service in late May. This followed a maintenance work which was intended to address a number of performance problems, but unfortunately encountered a hardware issue with one of the main network switches Comet relies on for interconnections between login, storage and (some) compute nodes. This took a very long time to resolve, due to a combination of sourcing the replacement parts under warranty, as well as arranging the specialist engineering staff to replace the faulty module and reinstall the replacement. Discussions are already under way between the various involved parties to minimise the possibility of such an issue happening again. Fortunately the other work which was planned during the maintenance does indeed appear to have been entirely successful: * Reconfiguration of the Comet uplinks to the main University campus network. This has **increased transfer speeds to/from RDW** as well as **reduced the latency/lag** in interactive SSH sessions for remote users. * Lustre file servers have been fitted with an additional 128GB of RAM in each node (2x metadata nodes, 2x object/data nodes) to provide an increased disk cache and to **smooth out high load situations**. * Lustre storage client has been downgraded on the two login nodes to match the client version installed on all storage and compute nodes. * NFS server has been configured to use a larger network buffer for incoming connections to **reduce dropped packets** in high load situations. * Minor changes to Open OnDemand partition options - removed legacy partitions (low latency, high memory) which are no longer selectable for interactive sessions. * All GPU nodes now back in operation To get through the backlog of compute work which built up throughout the extended downtime, the decision was made to make a number of changes to increase compute resources for the majority of users: * **Increase** the ratio of physical compute nodes allocated to free partitions (more standard nodes, one more GPU node) * **Increase** per-project resource limits for free and paid projects (task array limits are now larger, and simultaneous CPU cores and GPU cards are now larger for all types of projects - you can use [[https://hpc.researchcomputing.ncl.ac.uk/dokuwiki/doku.php?id=started:simple_slurm_tools|sproject]] from our Simple Slurm Tools software to view these) ---- [[:status:newsletters|Back to HPC Newsletters]]