MDC Downtime Starting 2026-08-14 for Emmy Phase 4 Installation
Starting early on the 14th of August 2026, the whole Modular Data Center (MDC) will be taken offline for essential maintenance and construction required for the upcoming cluster island Emmy Phase 4.
What will happen and how will it affect users?
During this transition:
- Most of Emmy Phase 2, including the
standard96,large96partitions and all variants, will be permanently retired, disassembled and removed. Some hardware in thelarge96partition will be kept, but will not be available during construction. - The legacy SCC
mediumpartition will be permanently decommissioned. - The legacy SCC login nodes
gwdu101-102aka.login-mdc.hpc.gwdg.dewill be retired (but the hardware will be repurposed under a different name). - JupyterHPC will operate with heavily reduced capacity, as most of the V100 and RTX5000 nodes will be decommissioned and other nodes will be unavailable during the downtime.
- The
lustre-mdcfilesystem will be inaccessible. - Limited access to the Future Technology Platform.
- The power and cooling infrastructure of the MDC will be fully overhauled, refurbished and expanded to accommodate the new hardware.
- New hardware will be installed, tuned and benchmarked.
The retirement of Phase 2 and the legacy SCC medium partition will significantly reduce the available compute hardware and will lead to much higher load on the remaining Emmy Phase 3 nodes in the medium/standard/large96s and scc-cpu partitions.
We expect a severe increase in queue sizes and thus waiting times before jobs will start until the new hardware comes online.
Why will it happen?
This downtime is necessary to clear physical space and prepare the infrastructure for the deployment of Emmy Phase 4. We recognize that working with constrained HPC resources will be challenging over the coming weeks and to some degree even months. Routine workflows and processes will be interrupted and more difficult than usual, and getting resources will take much longer. We apologize for the inconvenience this necessary transition will cause and thank you for your patience, flexibility, and continued understanding as we complete these upgrades.
How will GPU users be affected?
If you predominantly use GPU nodes (grete, react, scc-gpu, grete:interactive, …), this will not affect you directly, since most GPU hardware is located in our other datacenter, the RZG (Rechenzentrum Göttingen).
However, the reduced availability of CPU capacity might bring affected users to focus more on GPU-related work or try out alternative, GPU-based workflows for their existing work.
Older GPU nodes with NVIDIA V100 and RTX5000 (Turing) hardware, currently used for JupyterHPC, will also be retired.
If you are using interactive jobs on the jupyter partition, you will notice the reduced capacity, higher wait times and slower execution of workloads due to overprovisioning/sharing of resources in this partition.
How long will it take?
Please understand that for any large project with multiple contractors, vendors and partners involved, there is a large degree of uncertainty as any delay from one of the many interdependent but moving parts will change the whole schedule. As such, it is hard to make concrete predictions and even harder to make promises, so any dates we can announce here are to be taken with a grain of salt.
The major construction work and overhaul of the power infrastructure is expected to be complete by the end of August.
After that point, lustre-mdc and the remaining large96 nodes will be available again and JupyterHPC will regain a small portion of its lost capacity.
The installation of the new Emmy Phase 4 hardware, new network and cooling infrastructure will likely take another two to three weeks.
When the hardware is operational, it needs to be fully tested, the operating system, drivers and software stack adapted, performance tuned and benchmarked and any major bugs have to be sorted out before it can be opened to users.
Powerusers will be able to get early access, get a head-start in porting their workloads to the new system and help us find remaining issues.
It will likely be mid-October until the new hardware is available and open for NHR users.
For the SCC, we will put the nodes from the old medium partition back into operation as part of scc-cpu to bridge the gap until we get new hardware for HPC.NDS and the SCC, which is planned for next year.
As always, we encourage all SCC users with more demanding workloads to apply for full NHR compute projects.
Looking Ahead
While this maintenance period will demand a lot of endurance and patience from all of us, the final outcome will eventually be a more powerful and significantly more efficient computing environment. Emmy Phase 4 will deliver next-generation hardware with substantially higher performance, less energy consumption, and expanded resources for all your present and future research workloads.
This cluster island will initially comprise:
- 252 nodes with 256 cores of AMD EPYC 9745 “Turin” (Zen5) CPUs with 768 GiB of DDR5 memory and 3.8 TiB very fast local NVMe storage
- 22 turbo nodes with very high clock rates optimized for single-threaded or low thread-count jobs, 64-core AMD EPYC 9335 “Turin” (Zen5) CPUs, 768 GiB of DDR5 memory and 7.6 TiB local NVMe storage
- 6 visualization nodes for JupyterHPC, each with 4x NVIDIA RTX PRO 4000 Blackwell GPUs, 64-core AMD EPYC 9335 “Turin” (Zen5) CPUs, 768 GiB of DDR5 memory and 3.8 TiB local NVMe storage
- 3 new login nodes with 64 cores of AMD EPYC 9335 “Turin” (Zen5) CPUs, 768 GiB of DDR5 memory and 7.6 TiB local NVMe storage
With an extension next year of:
- 10 nodes with 256-core AMD “Venice” (Zen6) CPUs, 1 TiB of DDR5 memory, 3.8 TiB local NVMe storage
- 12 “large” nodes with 256-core AMD “Venice” (Zen6) CPUs, 4 TiB of DDR5 memory, 3.8 TiB local NVMe storage
- 2 GPU nodes with 8x NVIDIA B200 180 GiB VRAM, 256-core AMD “Venice” (Zen6) CPUs, 2 TiB of DDR5 memory, 7.6 TiB local NVMe storage
- 2 more login nodes with 256 cores AMD “Venice” (Zen6) CPUs, 1 TiB of DDR5 memory, 7.6 TiB local NVMe storage
All new nodes will be equipped with Cornelis CN5000 OmniPath adapters, with 200G connections and the new login and transfer nodes will additionally feature 2x25G (potentially 4x25G) Ethernet connections.