An earlier version of this page said pure MPI won at every size. A larger experiment says the opposite.
What I claimed
A single-node sweep over every legal rank/thread split of 128 cores.
| Layout (128 cores) | 1.00 M unknowns | 40.96 M unknowns |
|---|---|---|
128 ranks × 1 thread | 3.9 s | 846.9 s |
64 ranks × 2 threads | 4.4 s | 850.8 s |
32 ranks × 4 threads | 6.4 s | 866.4 s |
16 ranks × 8 threads | 15.5 s | 1140.0 s |
1 rank × 128 threads | 208.4 s | — |
// superseded
Pure MPI wins every column. I wrote that up as a finding.
What broke it
Two audits, neither about performance.
- The preconditioner changes with the rank count, so the rows were different algorithms.
- 302 of 303 logs had a process count that disagreed with their filename.
The experiment that replaced it
Rebuilt: CG + GAMG fixed everywhere, 5M–165M unknowns, up to 4,096 cores.
408 runs passed a six-point content audit; failures are kept but never aggregated.
The opposite answer
Hybrid wins in both dimensions, with a different layout each time.
| Experiment | Fastest layout | Result |
|---|---|---|
| 2D weak scaling | 64 × 2 | 794.8M eq/s @ 165M / 32 nodes · +13.1% vs flat MPI |
| 3D weak scaling | 32 × 4 | +33.2% throughput at best point · 24.9% time saved |
| 2D strong scaling, fixed 20M | 64 × 2 | 29.83× on 32 nodes · 93.2% efficiency |
| 3D strong scaling, fixed 20M | 16 × 8 | 3.81× faster than flat MPI on 32 nodes |
There is no single best thread count — only a best one for a given work per core.
Why, this time with a mechanism
MPI exposure collapses as OpenMP waiting rises; the ranking is where they cross.
| Layout | PETSc kernels | OpenMP waiting | MPI |
|---|---|---|---|
64 × 2 | 4.5% | 48.3% | 46.7% |
32 × 4 | 7.5% | 69.3% | 22.6% |
16 × 8 | 15.6% | 76.5% | 6.2% |
Threaded BLAS is not the explanation either.
What I actually changed
- Verify from the log body, never the filename.
- Hold the algorithm fixed before comparing hardware configurations.
- Time a steady state, not a first run.
- Publish the failures next to the results.