Skip to content

Emulators (surrogate models)

EmulatorTrainer fits a surrogate of the model's scalar observable features, and EmulatorBundle is the saved artefact plus the checks that decide whether it may be used. See Emulators for the workflow.

libcuflynx.emulators.emulator_trainer.EmulatorTrainer

EmulatorTrainer(param_id, emulator_settings, comm=None)

Design -> simulate -> fit -> validate -> persist, for one param-id engine.

Parameters:

Name Type Description Default
param_id

an ParamID built against the truth solver. Passing one whose sim helper is itself an emulator is rejected: it would fit a surrogate of a surrogate.

required
emulator_settings

the emulator_settings block (see ANALYSIS_OPTIONS).

required
comm

an MPI communicator; None means "discover one, or run serially".

None

There is deliberately no DEBUG shrinking here, unlike the optimiser and MCMC options. num_train_samples is the one setting that decides whether the emulator is any good, and quietly cutting it would produce an emulator that answers everything and is wrong -- the failure mode this whole feature is built to prevent. A cheap run asks for it by name.

feature_labels property

feature_labels

The emulator's outputs, named exactly as the run that will use it names them.

init_from_dict classmethod

init_from_dict(inp_data_dict, comm=None)

Build the truth-solver engine this config describes, then a trainer over it.

design

design(n_samples=None, sample_type=None, seed_offset=0)

Training points over the params_for_id box, shape (n_samples, num_params).

Deterministic given the seed, so every rank builds the identical design and no broadcast is needed to agree on who evaluates which sample. seed_offset separates one stage from the next: two stages of the same method on the same seed would otherwise draw the same points twice.

evaluate

evaluate(design)

Run the truth model at every design point; returns (x, y) on rank 0.

Contiguous block split across ranks then a gather, the same shape as the Sobol sampler's parallel loop. A sample whose simulation fails is dropped rather than imputed: an imputed training target is a fabricated observation, and the emulator would learn it as fact.

fit

fit(x, y, space_filling=None)

Fit and compare emulators; returns (model, r2, rmse, name, x_scale, y_scale).

Both x and y are mapped onto a well-conditioned range first. CA parameters routinely span a compliance near 1e-9 and a resistance near 1e8, and autoemulate works in float32 torch, where that spread alone is enough to ruin a kernel fit.

A test split is held out here, before fitting, and the reported R2/RMSE are scored on it. Scoring on the training points instead would report a number that says how well the emulator memorised the design, which is precisely the reassurance a bad emulator would give.

train

train()

The whole pipeline. Returns the bundle on rank 0, None elsewhere.

With reuse_samples set, the design and the simulations are skipped entirely and the samples a previous run saved are fitted instead -- see :meth:load_previous_samples. The rest of the path (fit, metadata, bundle, save) is the same one a fresh run takes, so the artefact is not a lesser kind of emulator.

libcuflynx.emulators.emulator_trainer.emulator_model_names

emulator_model_names()

The emulator names emulator_settings.models accepts, for a settings form.

Discovered from autoemulate's registry rather than hardcoded, in the same spirit as cost_func_metadata(); empty when autoemulate is not installed, so a tool can show the setting as unavailable instead of showing a stale list.

libcuflynx.emulators.emulator_trainer.resolve_emulator_dir

resolve_emulator_dir(inp_data_dict)

Where this config's emulator lives, defaulting beside the param-id output.

libcuflynx.emulators.emulator_bundle.EmulatorBundle

EmulatorBundle(
    model, meta, x_train=None, y_train=None, validation=None
)

A fitted emulator plus the metadata that makes it checkable.

Parameters:

Name Type Description Default
model

the fitted emulator (an autoemulate Emulator, or anything with a predict(x) returning an array-like of shape (n, n_features)).

required
meta

the metadata dict (see REQUIRED_META_KEYS).

required
x_train / y_train

the training design and targets, in real units. Kept so the emulator can be refitted, extended or audited without re-running the simulator.

required
validation

the held-out report from :func:emulators.emulator_trainer._validation_report -- per-feature statistics plus the points they were computed from.

None

predict

predict(theta, out_of_bounds='error')

Predicted scalar features for one theta (1-D) or many (2-D).

check_quality

check_quality(min_r2)

Refuse an emulator whose worst held-out R2 is below min_r2.

Named per feature, because "the emulator is bad" is not actionable while "max of aortic_root/u has R2 0.42" tells the user which observable to add samples for.

check_bounds

check_bounds(theta, policy='error')

Apply the out-of-training-box policy, returning the theta to evaluate.

An emulator is an interpolant; outside its design it is an extrapolation with no error estimate at all. Refusing is the default for that reason.

check_matches

check_matches(
    live_fingerprint,
    param_entry_labels=None,
    feature_labels=None,
)

Refuse a bundle trained against a different model, parameter set or protocol.

save

save(directory)

Write model, metadata and training data to directory. Returns the directory.

load classmethod

load(directory)

make_scale staticmethod

make_scale(values)

Shift/span for an affine map onto a well-conditioned range.

A span of zero (a constant column -- a parameter pinned to one value, or a feature the model does not respond to) would divide by zero, so it becomes 1: the column maps to a constant, which is exactly what it is.

libcuflynx.emulators.emulator_bundle.fingerprint

fingerprint(
    param_id_info, obs_info, protocol_info, model_path=None
)

A stable digest of everything an emulator was trained against.

Changing a parameter's bounds, adding a data_item, editing an operation, moving a sub-experiment or regenerating the model all change what theta -> features means. None of those changes make the old emulator raise on its own -- it would keep answering, about a different model. Comparing this digest is what turns that into an error.

model_path is hashed by content when it exists, so a regenerated CellML invalidates the emulator even if the inputs that produced it are unchanged.

The emulator is used through the ordinary simulation-helper interface, so every analysis reaches it the same way it reaches a solver.

libcuflynx.solver_wrappers.emulator_solver_helper.SimulationHelper

SimulationHelper(
    emulator_dir,
    dt=0.01,
    sim_time=1.0,
    solver_info=None,
    pre_time=0.0,
    bundle=None,
    out_of_bounds=None,
)

Emulator-backed drop-in for the solver helpers.

Parameters:

Name Type Description Default
emulator_dir

directory holding the trained bundle (emulator_metadata.json etc).

required
dt

output sampling step, kept only so the time bookkeeping matches the real backends.

0.01
sim_time

logged duration.

1.0
solver_info

accepted and ignored; there is no integrator to configure.

None
pre_time

unlogged spin-up; absorbed into the emulator at training time.

0.0
bundle

an already-loaded bundle, used instead of reading emulator_dir (tests, and callers that have validated one already).

None
out_of_bounds

'error' | 'warn' | 'clip' -- what to do off the training box.

None

set_theta

set_theta(theta)

Set the calibration vector directly, in params_for_id entry order.

This, not set_param_vals, is the emulator's real input. A modifier entry occupies one slot in theta but names several model parameters, and by the time the executor calls set_param_vals those slots have already been expanded to per-parameter values -- which is the wrong thing to feed a surrogate trained on theta. CA's call sites set theta here before the protocol runs, so no inversion is ever needed.

set_param_vals

set_param_vals(param_names, param_vals, change_states=True)

Accept the executor's per-parameter values.

With no modifiers in play these are theta itself, entry for entry, so they are recorded as such and the helper works for a caller that only knows the ordinary interface. With modifiers they are the expanded per-target values, which theta cannot be recovered from here -- set_theta must have been called, and this is then a consistency check.

run

run()

Predict the feature vector for the current theta. Returns success, like a solver.

get_results

get_results(variables_list_of_lists, flatten=False)

The predicted features, in the shape the executor expects from a solver.

Each operand slot holds a length-1 array carrying its data_item's predicted feature. The consumers that know about emulates_features read the value and skip the operation; anything else applying mean/max/min to it gets the same number back, which is the least surprising thing an unaware caller could receive.

get_predicted_features

get_predicted_features()

The predicted scalar feature per data_item, nan for any the emulator misses.

Indexed by data_item so both reduction sites can read it positionally.

set_obs_map

set_obs_map(const_idx_to_obs_idx, num_obs=None)

Tell the helper which data_item each trained feature belongs to.

The emulator's outputs are ordered by obs_info['const_idx_to_obs_idx'] while the consumers index by data_item. Rather than assume the two coincide, the caller states the mapping once at setup; without it they are taken to be the same, which is true whenever every data_item is a scalar one -- the only case the emulator supports.

update_times

update_times(dt, start_time, sim_time, pre_time)

Time bookkeeping only -- nothing is integrated, but tSim must still be right.

The protocol executor concatenates each sub-experiment's tSim onto a cumulative time vector and drops the duplicated first sample, so a missing or single-point axis would break the caller rather than the emulator. Mirrors the opencor backend.

get_init_param_vals

get_init_param_vals(param_names)

Parameter defaults, served from the snapshot taken when the emulator was trained.

The optimiser's x0 and resolve_modifier_baselines both read these before anything is simulated. The emulator has no model to read them from, so training recorded them.