Emulators (surrogate models)
EmulatorTrainer fits a surrogate of the model's scalar observable features, and
EmulatorBundle is the saved artefact plus the checks that decide whether it may be
used. See Emulators for the workflow.
libcuflynx.emulators.emulator_trainer.EmulatorTrainer
EmulatorTrainer(param_id, emulator_settings, comm=None)
Design -> simulate -> fit -> validate -> persist, for one param-id engine.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
param_id
|
an |
required | |
emulator_settings
|
the |
required | |
comm
|
an MPI communicator; |
None
|
There is deliberately no DEBUG shrinking here, unlike the optimiser and MCMC options.
num_train_samples is the one setting that decides whether the emulator is any good,
and quietly cutting it would produce an emulator that answers everything and is wrong --
the failure mode this whole feature is built to prevent. A cheap run asks for it by name.
feature_labels
property
feature_labels
The emulator's outputs, named exactly as the run that will use it names them.
init_from_dict
classmethod
init_from_dict(inp_data_dict, comm=None)
Build the truth-solver engine this config describes, then a trainer over it.
design
design(n_samples=None, sample_type=None, seed_offset=0)
Training points over the params_for_id box, shape (n_samples, num_params).
Deterministic given the seed, so every rank builds the identical design and no
broadcast is needed to agree on who evaluates which sample. seed_offset
separates one stage from the next: two stages of the same method on the same seed
would otherwise draw the same points twice.
evaluate
evaluate(design)
Run the truth model at every design point; returns (x, y) on rank 0.
Contiguous block split across ranks then a gather, the same shape as the Sobol sampler's parallel loop. A sample whose simulation fails is dropped rather than imputed: an imputed training target is a fabricated observation, and the emulator would learn it as fact.
fit
fit(x, y, space_filling=None)
Fit and compare emulators; returns (model, r2, rmse, name, x_scale, y_scale).
Both x and y are mapped onto a well-conditioned range first. CA parameters routinely span a compliance near 1e-9 and a resistance near 1e8, and autoemulate works in float32 torch, where that spread alone is enough to ruin a kernel fit.
A test split is held out here, before fitting, and the reported R2/RMSE are scored on it. Scoring on the training points instead would report a number that says how well the emulator memorised the design, which is precisely the reassurance a bad emulator would give.
train
train()
The whole pipeline. Returns the bundle on rank 0, None elsewhere.
With reuse_samples set, the design and the simulations are skipped entirely and the
samples a previous run saved are fitted instead -- see :meth:load_previous_samples.
The rest of the path (fit, metadata, bundle, save) is the same one a fresh run takes,
so the artefact is not a lesser kind of emulator.
libcuflynx.emulators.emulator_trainer.emulator_model_names
emulator_model_names()
The emulator names emulator_settings.models accepts, for a settings form.
Discovered from autoemulate's registry rather than hardcoded, in the same spirit as
cost_func_metadata(); empty when autoemulate is not installed, so a tool can show the
setting as unavailable instead of showing a stale list.
libcuflynx.emulators.emulator_trainer.resolve_emulator_dir
resolve_emulator_dir(inp_data_dict)
Where this config's emulator lives, defaulting beside the param-id output.
libcuflynx.emulators.emulator_bundle.EmulatorBundle
EmulatorBundle(
model, meta, x_train=None, y_train=None, validation=None
)
A fitted emulator plus the metadata that makes it checkable.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
the fitted emulator (an autoemulate |
required | |
meta
|
the metadata dict (see |
required | |
x_train / y_train
|
the training design and targets, in real units. Kept so the emulator can be refitted, extended or audited without re-running the simulator. |
required | |
validation
|
the held-out report from :func: |
None
|
predict
predict(theta, out_of_bounds='error')
Predicted scalar features for one theta (1-D) or many (2-D).
check_quality
check_quality(min_r2)
Refuse an emulator whose worst held-out R2 is below min_r2.
Named per feature, because "the emulator is bad" is not actionable while "max of aortic_root/u has R2 0.42" tells the user which observable to add samples for.
check_bounds
check_bounds(theta, policy='error')
Apply the out-of-training-box policy, returning the theta to evaluate.
An emulator is an interpolant; outside its design it is an extrapolation with no error estimate at all. Refusing is the default for that reason.
check_matches
check_matches(
live_fingerprint,
param_entry_labels=None,
feature_labels=None,
)
Refuse a bundle trained against a different model, parameter set or protocol.
save
save(directory)
Write model, metadata and training data to directory. Returns the directory.
load
classmethod
load(directory)
make_scale
staticmethod
make_scale(values)
Shift/span for an affine map onto a well-conditioned range.
A span of zero (a constant column -- a parameter pinned to one value, or a feature the model does not respond to) would divide by zero, so it becomes 1: the column maps to a constant, which is exactly what it is.
libcuflynx.emulators.emulator_bundle.fingerprint
fingerprint(
param_id_info, obs_info, protocol_info, model_path=None
)
A stable digest of everything an emulator was trained against.
Changing a parameter's bounds, adding a data_item, editing an operation, moving a sub-experiment or regenerating the model all change what theta -> features means. None of those changes make the old emulator raise on its own -- it would keep answering, about a different model. Comparing this digest is what turns that into an error.
model_path is hashed by content when it exists, so a regenerated CellML invalidates the
emulator even if the inputs that produced it are unchanged.
The emulator is used through the ordinary simulation-helper interface, so every analysis reaches it the same way it reaches a solver.
libcuflynx.solver_wrappers.emulator_solver_helper.SimulationHelper
SimulationHelper(
emulator_dir,
dt=0.01,
sim_time=1.0,
solver_info=None,
pre_time=0.0,
bundle=None,
out_of_bounds=None,
)
Emulator-backed drop-in for the solver helpers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
emulator_dir
|
directory holding the trained bundle ( |
required | |
dt
|
output sampling step, kept only so the time bookkeeping matches the real backends. |
0.01
|
|
sim_time
|
logged duration. |
1.0
|
|
solver_info
|
accepted and ignored; there is no integrator to configure. |
None
|
|
pre_time
|
unlogged spin-up; absorbed into the emulator at training time. |
0.0
|
|
bundle
|
an already-loaded bundle, used instead of reading |
None
|
|
out_of_bounds
|
'error' | 'warn' | 'clip' -- what to do off the training box. |
None
|
set_theta
set_theta(theta)
Set the calibration vector directly, in params_for_id entry order.
This, not set_param_vals, is the emulator's real input. A modifier entry occupies
one slot in theta but names several model parameters, and by the time the executor calls
set_param_vals those slots have already been expanded to per-parameter values --
which is the wrong thing to feed a surrogate trained on theta. CA's call sites set theta
here before the protocol runs, so no inversion is ever needed.
set_param_vals
set_param_vals(param_names, param_vals, change_states=True)
Accept the executor's per-parameter values.
With no modifiers in play these are theta itself, entry for entry, so they are recorded
as such and the helper works for a caller that only knows the ordinary interface. With
modifiers they are the expanded per-target values, which theta cannot be recovered from
here -- set_theta must have been called, and this is then a consistency check.
run
run()
Predict the feature vector for the current theta. Returns success, like a solver.
get_results
get_results(variables_list_of_lists, flatten=False)
The predicted features, in the shape the executor expects from a solver.
Each operand slot holds a length-1 array carrying its data_item's predicted feature.
The consumers that know about emulates_features read the value and skip the
operation; anything else applying mean/max/min to it gets the same number
back, which is the least surprising thing an unaware caller could receive.
get_predicted_features
get_predicted_features()
The predicted scalar feature per data_item, nan for any the emulator misses.
Indexed by data_item so both reduction sites can read it positionally.
set_obs_map
set_obs_map(const_idx_to_obs_idx, num_obs=None)
Tell the helper which data_item each trained feature belongs to.
The emulator's outputs are ordered by obs_info['const_idx_to_obs_idx'] while the
consumers index by data_item. Rather than assume the two coincide, the caller states
the mapping once at setup; without it they are taken to be the same, which is true
whenever every data_item is a scalar one -- the only case the emulator supports.
update_times
update_times(dt, start_time, sim_time, pre_time)
Time bookkeeping only -- nothing is integrated, but tSim must still be right.
The protocol executor concatenates each sub-experiment's tSim onto a cumulative
time vector and drops the duplicated first sample, so a missing or single-point axis
would break the caller rather than the emulator. Mirrors the opencor backend.
get_init_param_vals
get_init_param_vals(param_names)
Parameter defaults, served from the snapshot taken when the emulator was trained.
The optimiser's x0 and resolve_modifier_baselines both read these before anything
is simulated. The emulator has no model to read them from, so training recorded them.