CPMD is an MPI program: one SCF is spread over the ranks of a communicator, and one of those ranks, the parent, owns input and output. Embedding it in a host that is itself started under ``mpirun`` raises three questions: which ranks form one CPMD calculation, which rank's result counts, and who ends MPI. This page explains the answers ``cpmdc`` gives and why. Every rank is a copy of the host ================================ Under ``mpirun``, each rank runs the whole host program, and each rank's copy of ``libcpmdc`` calls into its own copy of CPMD. CPMD's collective operations tie those copies together: an SCF can only advance when every rank of its communicator takes part. So every rank must make the same ``cpmdc`` calls, with the same messages, in the same order. A rank that skips a call, or evaluates a different geometry, leaves the others waiting inside CPMD. The communicator CPMD uses ========================== OpenCPMD keeps its world communicator in ``mp_comm_world``. ``mp_start`` assigns that variable only when CPMD itself calls ``MPI_Init``. When MPI is already running, ``mp_start`` leaves ``mp_comm_world`` as the host set it. ``cpmdc_bind_calculator(ranks_per_calc)`` initialises MPI when it is not running, then splits ``MPI_COMM_WORLD`` with ``MPI_Comm_split``, colour ``world_rank / ranks_per_calc`` and key ``world_rank mod ranks_per_calc``, and installs the new communicator as ``mp_comm_world`` before ``mp_start``. Each colour is one calculator: its own CPMD setup, its own SCF, its own stored orbitals, all driven by the same deck. The split happens once per process. ``embed_calculator_bound``, a module variable in ``cpmd_embed_c_api``, records it, and a second call only reports the group index. ========== ================== =================== ============ World size ``ranks_per_calc`` Calculators Parent ranks ========== ================== =================== ============ 16 16, 0, or no call 1 0 16 4 4 0, 4, 8, 12 16 5 refused, returns -1 ========== ================== =================== ============ Why the parent rank's result counts =================================== CPMD computes the energy and ionic forces for the calculation as a whole, and its parent rank, rank 0 of ``mp_comm_world``, is the one CPMD itself treats as holding them: it writes the output, the restart file, and ``GEOMETRY``. The other ranks return from the same call with values of their own. They agreed with the parent in a 4-rank water single point. Over the many steps of a search they need not: in a 4-rank dimer search where every rank also ran the host's optimiser, the ranks' paths diverged, and only rank 0's results matched the file route. The robust rule is that the host takes the parent's result and broadcasts it, so every rank's optimiser steps from the same numbers. ``cpmdc`` does not broadcast on its own because it cannot know which of the host's communicators should receive the result. One session per calculator ========================== CPMD's module state and the stored orbitals live in the process, one set per rank. A calculator therefore evaluates one ``CPMDCSession``. A second session in the same process is possible, but switching between them re-applies the configuration and clears the stored orbitals each time, which turns every call into a cold start. Independent calculations belong on separate calculators: one per image of a nudged elastic band, one per end of a dimer. Who ends MPI ============ CPMD calls ``MPI_Init`` in its setup when MPI is not running yet, and ``cpmdc_bind_calculator`` does the same. Neither ``cpmdc`` nor the patched CPMD calls ``MPI_Finalize``: in a host, the host decides when MPI ends, and a library that finalized MPI would break a host that still needs it. The host has to finalize on every rank before exit. Open MPI 5 counts a rank that exits without it as an abnormal termination and kills the others, which in a 4-rank single point killed rank 0 before it had written its result. ``cpmdc_finalize()`` marks the library finalized and leaves MPI alone for the same reason. Threads inside ranks ==================== An OpenCPMD build with OpenMP adds threads inside each rank. Ranks, threads, and calculators multiply: 16 cores can hold one calculator of 16 single-threaded ranks, four calculators of four ranks, or other combinations, set with ``mpirun -np``, ``cpmdc_bind_calculator``, and ``OMP_NUM_THREADS``. The :doc:`mpirun how-to <../howto/mpi>` shows the calls.