The Quantum Exact Simulation Toolkit v4.3.0
Loading...
Searching...
No Matches
Experimental

Experimental functions with tentative APIs. More...

Functions

Qureg createQuregFromFile (const char *fn)
 
int getQuESTNumGpuThreadsPerBlock ()
 
void initCustomMpiCommQuESTEnv (MPI_Comm questComm, int useGpuAccel, int useMultithread)
 
void initCustomMpiQuESTEnv (int useDistrib, bool userOwnsMpi, int useGpuAccel, int useMultithread)
 
void saveQuregToFile (Qureg qureg, const char *fn)
 
void setQuESTNumGpuThreadsPerBlock (int numThreadsPerBlock)
 

Detailed Description

Experimental functions with tentative APIs.

Function Documentation

◆ createQuregFromFile()

Qureg createQuregFromFile ( const char * fn)

Creates a new Qureg from a file (or folder) previously created by saveQuregToFile(), with automatically chosen deployments (independent of those used when the file was saved), and populates the Qureg with the saved amplitudes.

The chosen deployments are identical to those chosen by createQureg() and createDensityQureg().

Note
The number of distributed nodes chosen by the autodeployer must agree with the number of nodes of the originally saved Qureg, else a error is thrown. Therefore, the number of MPI processes calling these functions cannot be changed between saveQuregToFile() and createQuregFromFile(), unless the Qureg was non-distributed in both settings.
Important
This function is only callable when QuEST is compiled with CMake option QUEST_ENABLE_ADIOS2=1.
Parameters
[in]fnthe file (or folder) path previously created by saveQuregToFile().
Returns
A new Qureg instance matching the saved dimension and amplitudes.
Exceptions
error
  • if QuEST was not compiled with CMake option QUEST_ENABLE_ADIOS2=1.
  • if fn cannot be read (since, for example, it does not exist).
  • if the precision of the saved Qureg differs from the current QuEST precision.
  • if the number of distributed nodes of the saved Qureg differs from the autodeployer's chosen number.
  • if the recorded Qureg dimensions would overflow the qindex type.
  • if the recorded total Qureg memory would overflow the size_t type.
  • if the system contains insufficient RAM (or VRAM) to store the Qureg in any deployment.
  • if any Qureg memory allocation unexpectedly fails.
See also
Author
Ashmit JaiSarita Gupta
Tyson Jones (input validation)

Definition at line 240 of file experimental.cpp.

240 {
241 validate_adios2IsCompiled(__func__);
242
243#if QUEST_COMPILE_ADIOS2
244
245 // pedantic but safe - don't let ADIOS2 start reading while other processes are working
246 if (comm_isActive())
247 comm_sync();
248
249 // make ADIOS2 MPI-aware even when the subsequently-loaded Qureg is
250 // auto-deployed to be non-distributed; every process will safely
251 // parse the file and independently update its Qureg copy
252 bool giveAdiosMpi = comm_isActive();
253
254 // gratuitously re-create ADIOS2 at every call, for simplicity (occluded by IO)
255 adios2::ADIOS adios = createAdios(giveAdiosMpi);
256 adios2::IO io = adios.DeclareIO("QuESTQuregLoad");
257
258 // use BP5 specifically to avoid non-root-hangs upon rank exceptions
259 io.SetEngine("BP5");
260
261 // attempt to open the file, and prepare to parse
262 adios2::Engine engine; // default ctor
263 bool success = false;
264 try {
265 engine = io.Open(fn, adios2::Mode::ReadRandomAccess);
266 success = true;
267 } catch (...) {}
268 validate_adiosCanOpenFileOnAllNodes(success, fn, __func__);
269
270 // check that the file contains the expected variables
271 auto vNumQubits = io.InquireVariable<int>("numQubits");
272 auto vNumNodes = io.InquireVariable<int>("numNodes");
273 auto vIsDensMatr = io.InquireVariable<int>("isDensityMatrix");
274 auto vQrealBytes = io.InquireVariable<size_t>("qrealBytes");
275 auto vAmpComponents = io.InquireVariable<qreal>("ampComponents");
276 bool areAllVarsPresent = vNumQubits && vNumNodes && vIsDensMatr && vQrealBytes && vAmpComponents;
277 validate_adiosFileContainsFieldsOnAllNodes(areAllVarsPresent, __func__);
278
279 // read dimension + precision metadata first, so we can size the new Qureg
280 int numQubits = 0;
281 int numNodes = 0;
282 int isDensMatr = 0;
283 size_t fileQrealBytes = 0;
284 success = false;
285 try {
286 engine.Get(vNumQubits, numQubits);
287 engine.Get(vNumNodes, numNodes);
288 engine.Get(vIsDensMatr, isDensMatr);
289 engine.Get(vQrealBytes, fileQrealBytes);
290 engine.PerformGets();
291 success = true;
292 } catch (...) {}
293 validate_adiosCanReadFileOnAllNodes(success, fn, __func__);
294
295 // check the amps are of the expected precision, and so are parsable
296 validate_newQuregFileMatchesPrecision(fileQrealBytes, __func__);
297
298 // attempt to create a matching-dimension Qureg with automatically chosen deployments
299 Qureg qureg = validateAndCreateCustomQureg(numQubits, isDensMatr,
300 modeflag::USE_AUTO, modeflag::USE_AUTO, modeflag::USE_AUTO, __func__);
301
302 // auto-distribution MUST match checkpointed distribution (pre-free to avoid leak)
303 if (qureg.numNodes != numNodes)
304 destroyQureg(qureg);
305 validate_newQuregNumNodesMatchesSavedFile(numNodes, qureg.numNodes, comm_getNumNodes(), numQubits, isDensMatr, __func__);
306
307 // read only this node's slice of the global amplitude array into its buffer
308 // (guaranteed not to overflow by above validateAndCreateCustomQureg validation)
309 qindex localReals = 2 * qureg.numAmpsPerNode;
310 qindex startReal = 2 * ((qindex) qureg.rank) * qureg.numAmpsPerNode;
311 vAmpComponents.SetSelection({ { (size_t) startReal }, { (size_t) localReals } });
312 success = false;
313 try {
314 engine.Get(vAmpComponents, reinterpret_cast<qreal*>(qureg.cpuAmps)); // immediate; PerformGets redundant
315 success = true;
316 } catch (...) {}
317 validate_adiosCanReadFileOnAllNodes(success, fn, __func__);
318
319 // complete ADIOS2 work
320 success = false;
321 try {
322 engine.Close();
323 success = true;
324 } catch (...) {}
325 validate_adiosCanReadFileOnAllNodes(success, fn, __func__);
326
327 // propagate the restored CPU amplitudes to the GPU, if deployed
328 if (qureg.isGpuAccelerated)
329 gpu_copyCpuToGpu(qureg);
330
331 return qureg;
332#else
333 // unreachable: the validation above always throws in non-checkpointing builds
334 return Qureg{};
335#endif
336}
void destroyQureg(Qureg qureg)
Definition qureg.cpp:340
Definition qureg.h:49

Referenced by createQuregFromFile().

◆ getQuESTNumGpuThreadsPerBlock()

int getQuESTNumGpuThreadsPerBlock ( )
Note
Documentation for this function or struct is under construction!
Author
Oliver Brown

Definition at line 137 of file experimental.cpp.

137 {
138 validate_envIsInit(__func__);
139
140 return gpu_getNumThreadsPerBlock();
141}

Referenced by TEST_CASE().

◆ initCustomMpiCommQuESTEnv()

void initCustomMpiCommQuESTEnv ( MPI_Comm questComm,
int useGpuAccel,
int useMultithread )
Note
Documentation for this function or struct is under construction!

Advanced initialiser which allows the user to provide an MPI communicator for QuEST to use. Use of this initialiser implies userOwnsMpi = true, (exposed by initCustomMpiQuESTEnv) and therefore that they have already initialised MPI, and they will call MPI_Finalize at the appropriate time.

The user-provided MPI communicator undergoes the same validation procedure as any that QuEST would use, and so must contain a power-of-2 number of processes.

Important
This function is only compiled and exposed when macro QUEST_COMPILE_SUBCOMM is 1, as is defined when providing CMake option QUEST_ENABLE_SUBCOMM during building.
Author
Oliver Brown

Definition at line 113 of file experimental.cpp.

113 {
114
115 // useDistrib and userOwnsMpi are implied by the user of this initialiser
116 const int useDistrib = 1;
117 const bool userOwnsMpi = true;
118
119 // pre-validate that we are able to set the MPI communicator
120 validate_mpiInitStatus(useDistrib, userOwnsMpi, __func__);
121 validate_mpiSubCommIsNonNull(userQuestComm != MPI_COMM_NULL, __func__);
122
123 // avoid re-setting the MPI comm (to avoid an internal error), which happens
124 // if a user illegally re-calls this function, which will be subsequently
125 // caught by the validation in validateAndInitCustomQuESTEnv() below
126 if (!comm_isActive()) {
127 bool success = comm_setMpiComm(userQuestComm, userOwnsMpi);
128 validate_mpiSubCommSetSucceeded(success, __func__);
129 }
130
131 // perform remaining validation (some is harmlessly repeated) and init QuEST env
132 validateAndInitCustomQuESTEnv(useDistrib, userOwnsMpi, useGpuAccel, useMultithread, __func__);
133}

◆ initCustomMpiQuESTEnv()

void initCustomMpiQuESTEnv ( int useDistrib,
bool userOwnsMpi,
int useGpuAccel,
int useMultithread )
Note
Documentation for this function or struct is under construction!

Advanced initialiser which lets the user positively declare that they take responsibility for MPI. This means we assume they have called MPI_Init, and that they will call MPI_Finalize.

Author
Oliver Brown

Definition at line 107 of file experimental.cpp.

107 {
108 validateAndInitCustomQuESTEnv(useDistrib, userOwnsMpi, useGpuAccel, useMultithread, __func__);
109}

◆ saveQuregToFile()

void saveQuregToFile ( Qureg qureg,
const char * fn )

Writes the contents of qureg to the file (or folder) fn, so that it may later be restored with createQuregFromFile(), potentially in another process.

The output records the qureg dimension (number of qubits and whether it is a density matrix), the amplitude precision, the Qureg's distribution, and the Qureg's full set of amplitudes. Other deployment information, such as whether the Qureg is multithreaded or GPU-accelerated, is not recorded.

There is no particular file extension or folder name suffix required, though since saving is performed with ADIOS2, a suffix of .bp is conventional.

Attention
Specifying fn equal to an existing directory or file will cause erasure and overwriting of its contents. It is especially dangerous to pass fn equal to a system directory, such as / on Unix, and may cause system corruption.
Important
This function is only callable when QuEST is compiled with CMake option QUEST_ENABLE_ADIOS2=1.
Parameters
[in]quregthe Qureg to write to disk.
[in]fnthe output file (or folder) path.
Exceptions
error
  • if qureg is uninitialised.
  • if QuEST was not compiled with CMake option QUEST_ENABLE_ADIOS2=1.
  • if opening or writing to fn fails.
See also
Author
Ashmit JaiSarita Gupta

Definition at line 155 of file experimental.cpp.

155 {
156 validate_adios2IsCompiled(__func__);
157 validate_quregFields(qureg, __func__);
158
159 (void) fn; // suppress unused warning
160
161#if QUEST_COMPILE_ADIOS2
162
163 // When the QuEST env is distributed, but the given Qureg is not (and is instead
164 // duplicated upon every node), we should permit only a single process (the root)
165 // to use ADIOS2 to write to the (assumably, shared) filesystem. Note that non-root
166 // nodes must not exit; they need to participate in validation syncs
167 bool shouldSkipAdios = (! qureg.isDistributed) && (comm_getRank() > ROOT_RANK);
168
169 // pedantic but safe - don't let ADIOS2 start reading amps prematurely
170 if (qureg.isDistributed)
171 comm_sync();
172
173 // TODO:
174 // We can optimise in GPU settings by giving ADIOS2 the device memory
175 // pointers; but for now, we simply stage into CPU memory first
176 if (qureg.isGpuAccelerated)
177 gpu_copyGpuToCpu(qureg);
178
179 // gratuitously re-create ADIOS2 at every call, for simplicity (occluded by IO)
180 adios2::ADIOS adios = createAdios(qureg.isDistributed);
181 adios2::IO io = adios.DeclareIO("QuESTQuregSave");
182
183 // use BP5 specifically to avoid non-root-hangs upon rank exceptions
184 io.SetEngine("BP5");
185
186 // attempt to open the file
187 adios2::Engine engine; // default ctor
188 bool success = false;
189 try {
190 if (!shouldSkipAdios)
191 engine = io.Open(fn, adios2::Mode::Write);
192 success = true;
193 } catch (...) {}
194 validate_adiosCanOpenFileOnAllNodes(success, fn, __func__);
195
196 // global single-value metadata; we deliberately record only the dimension
197 // and precision, never incidental deployment fields (the loader chooses its
198 // own deployment) nor derivable fields (like numAmps)
199 adios2::Variable<int> vNumQubits = io.DefineVariable<int>("numQubits");
200 adios2::Variable<int> vNumNodes = io.DefineVariable<int>("numNodes");
201 adios2::Variable<int> vIsDensMatr = io.DefineVariable<int>("isDensityMatrix");
202 adios2::Variable<size_t> vQrealBytes = io.DefineVariable<size_t>("qrealBytes"); // also encodes precision
203
204 // amplitudes are stored as interleaved (real, imag) reals to stay agnostic
205 // to precision and to ADIOS2's complex-type support; each node writes only
206 // its local slice into the global array, avoiding excessive memory use
207 // (these scalars are guaranteed not to overflow by createQureg validation)
208 qindex globalReals = 2 * qureg.numAmps;
209 qindex localReals = 2 * qureg.numAmpsPerNode;
210 qindex startReal = localReals * qureg.rank;
211 adios2::Variable<qreal> vAmpComponents = io.DefineVariable<qreal>(
212 "ampComponents",
213 { (size_t) globalReals },
214 { (size_t) startReal },
215 { (size_t) localReals });
216
217 // attempt to write to file
218 success = false;
219 try {
220 if (!shouldSkipAdios) {
221 engine.Put(vNumQubits, qureg.numQubits);
222 engine.Put(vNumNodes, qureg.numNodes);
223 engine.Put(vIsDensMatr, qureg.isDensityMatrix);
224 engine.Put(vQrealBytes, sizeof(qreal));
225 engine.Put(vAmpComponents, reinterpret_cast<qreal*>(qureg.cpuAmps));
226 engine.Close();
227 }
228 success = true;
229 } catch (...) {}
230 validate_adiosCanWriteToFileOnAllNodes(success, fn, __func__);
231
232 // prevent any process from continuing until ADIOS2 is fully finished
233 if (qureg.isDistributed)
234 comm_sync();
235
236#endif
237}

Referenced by saveQuregToFile().

◆ setQuESTNumGpuThreadsPerBlock()

void setQuESTNumGpuThreadsPerBlock ( int numThreadsPerBlock)

Overrides the number of CUDA threads per block (or blockDim) used by QuEST's GPU-accelerated backend.

This changes the GPU parallelisation granularity and can affect performance, and is useful for performance tuning or diagnostics. Before this function is called, QuEST will use the number as specified by the environment variable QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK, if defined. Otherwise, it will use the value specified by the CMake/compile option of the same name, which itself presently defaults to 128. After this function is called, QuEST will adopt numThreadsPerBlock for the remainder of execution, or until this function is called again.

Practical values of numThreadsPerBlock can vary with the simulation size, the user's GPU hardware, and whether it is NVIDIA or AMD, which have respective warp sizes of 32 and 64.

Note
This function has no effect when QuEST is not deployed with GPU-acceleration enabled.
Parameters
[in]numThreadsPerBlockthe new block size.
Exceptions
error
  • if the QuESTEnv is not initialised.
  • if numThreadsPerBlock is negative.
  • if numThreadsPerBlock is not a multiple of the GPU warp size.
  • if numThreadsPerBlock exceeds the maximum blockDim imposed by the GPU hardware.
See also
Author
Oliver Brown
Tyson Jones

Definition at line 144 of file experimental.cpp.

144 {
145 validate_envIsInit(__func__);
146
147 // validation messages and queries depend upon GPU usage
148 bool gpuIsActive = getQuESTEnv().isGpuAccelerated;
149 validate_numGpuThreadsPerBlock(numTPB, gpuIsActive, __func__);
150
151 gpu_setNumThreadsPerBlock(numTPB);
152}
QuESTEnv getQuESTEnv()

Referenced by TEST_CASE().