Goal¶
Run a host that embeds cpmdc on several MPI (Message Passing
Interface) ranks, so that CPMD parallelises each self-consistent field
(SCF) calculation over the ranks, and keep every rank’s view of the
result consistent. This page needs an OpenCPMD build of cpmdc; the
default build runs the reference evaluator on each rank independently.
The rules¶
Start every rank of the host under
mpirun. CPMD initialises MPI on the first evaluation when the host has not.Every rank makes the same
cpmdccalls with the same messages in the same order. An evaluation is collective over the ranks of its calculator.Take the energy and forces from the first rank of each calculator, CPMD’s parent rank, and broadcast them to the other ranks. The other ranks return from the call with values of their own. In a 4-rank water single point all four printed the parent’s energy and forces, but in a 4-rank dimer search, where every rank also ran the host’s optimizer, rank 1 took 3 steps on different curvatures while rank 0 took 5, and only rank 0’s results matched an independent
cpmd.xrun.Keep one
CPMDCSessionper calculator.Call
MPI_Finalizeon every rank before the host exits.cpmdc_finalize()does not finalize MPI.
One calculator over all ranks¶
Without further setup, CPMD uses MPI_COMM_WORLD, so all ranks form
one calculator:
export CPMDC_PSEUDO_DIR=/path/to/pseudopotentials
mpirun -np 16 ./host params.bin step.bin
Rank 0 of MPI_COMM_WORLD is then CPMD’s parent rank.
Several calculators in one job¶
cpmdc_bind_calculator(ranks_per_calc) splits MPI_COMM_WORLD into
groups of ranks_per_calc consecutive ranks. Each group becomes one
CPMD calculator with its own communicator and its own wavefunction,
running the same deck. Call it on every rank, once, before the first
evaluation:
#include <cpmdc.h>
#include <mpi.h>
int main(int argc, char **argv) {
MPI_Init(&argc, &argv);
int group = cpmdc_bind_calculator(4); /* 16 ranks -> 4 calculators */
if (group < 0) {
/* world size not a multiple of 4, or no CPMD backend */
MPI_Abort(MPI_COMM_WORLD, 1);
}
int world_rank = 0;
MPI_Comm_rank(MPI_COMM_WORLD, &world_rank);
int is_parent = (world_rank % 4) == 0;
/* group 0 evaluates image 0, group 1 image 1, ... */
/* ... create one session, evaluate, share the parent's result ... */
cpmdc_finalize();
MPI_Finalize();
return 0;
}
Argument or result |
Meaning |
|---|---|
|
ranks per calculator; the world size must be a multiple of it |
|
one calculator over the whole world |
return value |
index of this rank’s calculator,
|
return value |
the split was refused, or the library has no CPMD backend (the default build and the stub) |
A second call returns the same index without splitting again. The
group’s first rank, world rank group * ranks_per_calc, is its parent
rank. cpmdc remembers the split in embed_calculator_bound.
mp_start assigns mp_comm_world only when CPMD itself calls
MPI_Init, and cpmdc_bind_calculator has already initialised MPI
and stored the calculator communicator by the time setup reaches
mp_start.
One session per calculator¶
The OpenCPMD module state and the stored orbitals belong to the process,
not to a CPMDCSession. When a process evaluates a second session,
cpmdc applies that session’s configuration and clears the stored
orbitals; the next call of either session starts cold. Run independent
calculations, such as the images of a nudged elastic band, on separate
calculators, one session each.
Finalize MPI¶
CPMD calls MPI_Init during its setup, and nothing in cpmdc calls
MPI_Finalize. Open MPI 5 treats a rank that exits without
MPI_Finalize as an abnormal termination and kills the ranks still
running; mpirun then reports a rank “exiting improperly” and lists a
missing finalize among the reasons. Call MPI_Finalize on every rank,
after cpmdc_finalize(), as the example above does. When the host
cannot restructure its exit path, register a handler with atexit
that calls MPI_Finalize when MPI_Initialized reports true and
MPI_Finalized reports false.
Threads¶
CPMD also runs OpenMP threads inside each rank when OpenCPMD was built
with -omp. Set OMP_NUM_THREADS per rank so that ranks times
threads matches the cores the job holds.