If 1 CPU can run a program, then 8 CPUs should make it much faster.
Right?
Not necessarily.
This is one of the first things that can be surprising when learning HPC.
We can keep adding CPU cores to a system, but that does not mean an application will keep getting proportionally faster.
The reason comes down to parallelism.
Not every part of a program can run in parallel
Imagine an application that takes 100 seconds to run on one CPU.
Suppose:
- 80 seconds can be parallelised
- 20 seconds must run sequentially
The 80 second part can use multiple CPU cores.
But the 20 second part still has to run one step after another.
So even if we give the application 100 CPU cores, that sequential portion doesn't magically disappear.
This is the basic idea behind Amdahl's Law.
What happens as we add more cores?
Imagine the application running on:
1 core → 100 seconds
2 cores → 60 seconds
4 cores → 40 seconds
8 cores → 30 seconds
16 cores → 25 seconds
32 cores → 23 seconds
We're adding more and more CPU cores.
But the improvement gets smaller.
Why?
Because eventually the parts of the application that cannot be parallelised become the limiting factor.
The application has reached a point where adding more CPUs provides very little benefit.
The problem isn't always the CPU
This is an important idea in HPC.
When an application doesn't scale well, the CPU may not be the problem.
The application might be waiting for:
- Memory
- Network communication
- Storage I/O
- Synchronisation between processes
- Locks
- Other processes
For example, imagine an MPI application running across many compute nodes.
Each process may need to exchange data with other processes.
Adding more nodes gives you more CPUs, but it also means more communication between them.
At some point, the communication overhead can start becoming significant.
You have more computing resources, but you are spending more time coordinating them.
Strong scaling vs weak scaling
This leads to two important HPC concepts.
Strong scaling asks:
How much faster can I solve the same problem by adding more resources?
For example:
Same problem
↓
8 cores → 16 cores → 32 cores → 64 cores
Ideally, the runtime keeps decreasing.
But eventually the speedup usually becomes smaller.
Weak scaling asks something different:
Can I solve a larger problem by adding more resources while keeping the workload per resource roughly the same?
For example:
8 cores → small problem
16 cores → larger problem
32 cores → even larger problem
Both are important ways of understanding HPC application performance.
More cores can even make things worse
This can sound strange, but adding resources can sometimes reduce efficiency.
Imagine 16 processes constantly exchanging data.
Now increase that to 128 processes.
There are many more processes that may need to communicate and synchronise.
If the application's communication pattern doesn't scale well, the additional resources can create more overhead than useful computation.
The CPUs aren't necessarily doing more useful work.
They may be spending more time waiting.
This is why benchmarking matters
You cannot determine how well an HPC application will scale simply by looking at the number of CPU cores in a server.
You need to test it.
For example:
1 node → 100 seconds
2 nodes → 55 seconds
4 nodes → 32 seconds
8 nodes → 25 seconds
16 nodes → 23 seconds
The application is scaling.
But not linearly.
Doubling the resources does not always halve the runtime.
That is normal for many real world HPC applications.
The goal isn't always maximum CPU usage
Another important lesson is that 100% CPU utilisation doesn't automatically mean an application is scaling well.
An application could be using every CPU core while spending significant time waiting for memory, communicating with other processes or performing synchronisation.
This is why HPC performance analysis looks at much more than CPU utilisation.
You may need to investigate:
- CPU utilisation
- Memory bandwidth
- NUMA behaviour
- Network traffic
- MPI communication
- Storage I/O
- Process placement
- CPU affinity
The bottleneck can move as you add resources.
So how many CPU cores should you use?
There isn't one answer.
The right number depends on the application.
For some workloads:
8 cores → good scaling
16 cores → better
32 cores → better
64 cores → little additional benefit
For another application, 64 or 128 cores might still provide excellent scaling.
This is why HPC systems don't simply try to give every application as many CPUs as possible.
The goal is to use the resources efficiently.
The bigger HPC lesson
HPC isn't simply about having more CPUs.
It's about using parallel resources effectively.
A well designed application can divide a large problem across many CPU cores or nodes and achieve significant speedup.
But parallelism has limits.
Sequential work, communication, synchronisation, memory bandwidth and I/O can all become bottlenecks.
So the next time someone says:
"This server has twice as many CPU cores, so the application should be twice as fast."
The better question is:
"Does the application scale?"
That's one of the fundamental questions in HPC performance.
And sometimes, the answer is:
not as much as you might expect.
Top comments (0)