Running Evaluations
Run full episodes or drive a model and environment step by step.
Use run() to score a model over complete episodes. Use session() when you need to inspect or control each step. Both drive the same reset, predict, step loop and apply the adapter resolved from the environment’s contract and the model’s ModelSpec.
Pick an entry point
| Entry point | You get | Reach for it when |
|---|---|---|
run() |
One call drives whole episodes and returns a typed RunResult. |
Scoring a model: leaderboards, sweeps, CI checks. |
session() |
A Session you step by hand (reset / predict / step). |
Rendering, custom stop conditions, branching, or mixing your own per-step logic. |
Session.reader / Session.read |
A read-only, role-addressed view of each raw observation. | Inspecting an env, logging canonical roles, or shaping a reward. |
read and reader let you inspect an observation by role while using a session.
run()whole episodessession()step by handread() / reader()inspect by rolerun(): the automated rollout
run() completes the episodes and returns a RunResult:
result = model.run(env, seeds=range(100))
print(f"mean reward {result.mean_reward:.2f}")
if result.success_rate is not None:
print(f"{result.success_rate:.0%} success")
For code you will reuse or package, define an EnvFactory and a Model subclass, then call MyPolicy().run(MyEnv(), episodes=10). env also accepts an existing Gymnasium-style environment, a remote handle, or an address:
result = model.run("tcp://127.0.0.1:5555", seeds=range(100))
The module-level run() accepts a Model class or instance, a remote model, or rlmesh.RANDOM_SAMPLE for an action-space sampling baseline. Wrap a prediction function in the backend Model that matches its array type:
import rlmesh
import rlmesh.numpy
result = rlmesh.run(rlmesh.numpy.Model(my_policy_fn), env, seeds=range(10))
baseline = rlmesh.run(rlmesh.RANDOM_SAMPLE, env, episodes=10)
Arguments
| Argument | Default | Meaning |
|---|---|---|
seeds |
None |
One seed per episode; alone, its length sets the episode count. |
episodes |
None |
Exact episode count; must equal the length of seeds when both are given. 0 runs none. |
execution_horizon |
1 |
Actions executed per predicted chunk; see Performance and Scaling. |
close_env |
False |
Shut the env down when the run finishes (opt-in). |
trial_index_base |
0 |
First trial ordinal for environments that declare benchmark trials. |
With neither seeds nor episodes, run() does a single episode. execution_horizon is accepted by both the bound methods (model.run / model.session) and the module-level run() / session(), which forwards it through.
model.run() drives the native runtime loop. It accepts hooks= for step and episode callbacks and instruction= for a model with a declared text input. For live viewing, use session() with view=.
Watching and capping the loop
run() also takes max_episode_steps and max_episode_seconds, per-episode caps that mark a capped episode truncated like an env time limit. These caps need the runtime to own resets; use episodes to bound an autoresetting vector env.
Pass hooks= to model.run() or Session.run. A RunHooks subclass can observe on_run_start, on_episode_start, on_step, on_episode_end, and on_run_end. StepEvent includes the step’s observation, action, reward, terminal flags, timings, and a lazy role read. The callbacks run in the same per-episode order on both paths. A hook exception aborts the run, and on_run_end fires once with the completed episodes:
class Progress(rlmesh.RunHooks):
def on_episode_end(self, result):
print(f"episode {result.index}: reward {result.reward:.2f}")
result = model.run(env, seeds=range(50), max_episode_steps=500, hooks=Progress())
The result
RunResult is immutable and aggregates its episodes:
| Member | Type | Meaning |
|---|---|---|
.episodes |
tuple[EpisodeResult, ...] |
One EpisodeResult per episode. |
.mean_reward |
float |
Mean total reward across episodes. |
.success_rate |
float | None |
Fraction the env reported as successful; None if any outcome is unknown. |
.num_episodes |
int |
Episode count. |
.total_steps |
int |
Summed steps across episodes. |
Each EpisodeResult carries index, seed, steps, reward, terminated, truncated, success, and timing fields (duration_s, predict_ms, step_ms):
for ep in result.episodes:
print(ep.index, ep.seed, ep.steps, ep.reward, ep.terminated, ep.success)
On the native run path, RunResult.telemetry contains aggregated timing and
size measurements. result.format_telemetry() prints them as a table. A
Session.run() result has no runtime telemetry rows; its episode timing
fields are still available.
session(): manual, step-by-step control
session() hands back a Session you drive yourself. Use it as a context manager so the env connection (and any managed model) closes on exit:
with model.session(env, instruction="put the cup on the plate") as sess:
obs, info = sess.reset(seed=0)
while not sess.done:
action = sess.predict(obs)
obs, reward, terminated, truncated, info = sess.step(action)
The loop primitives mirror Gymnasium, with the adapter folded in.
reset(seed)predict(obs)step(action)
- not done: back to predict(obs)
- done: back to reset(seed)
sess.reset(seed=None, trial_index=None)→(obs, info). Begins an episode; ends the previous one (firingon_episode_end) and clears adapter state such as the frame-stack buffer.trial_indexis the 0-based ordinal of this episode in a benchmark’s trial sweep; it reaches the env asreset(options={"trial_index": ...}), but only if the env declared the key inEnvFactory.reset_options– passing one to an env that did not warns and resets without it.sess.run(trial_index_base=0)walks the ordinals for you (episodeiis trialtrial_index_base + i), delivers each to a declaring env, and reports each onEpisodeResult.trialwhether or not the env asked for it.sess.predict(obs)→action. Applies the model’s adapter around the model’s own predict: the declarative obs transform, host-side frame stacking, anyCustomcode, instruction injection into declared text leaves, and chunk replay (one action per call). Returns an env-ready action.sess.step(action)→(obs, reward, terminated, truncated, info). Applies the action and records reward and termination.sess.doneisTrueonce the current episode terminated or truncated.sess.close()releases the connection, shuts the env down only on theclose_envopt-in, and fireson_close.
Use session() when you need to branch on each step, render a frame, or stop early. The module-level session() accepts the same model forms as run(), including rlmesh.RANDOM_SAMPLE.
on_episode_end fires at every episode boundary (the next reset(), or close() for the last episode), so a stateful model clears its per-episode state on either path. Session.run pumps whole episodes through the session primitives; model.run() drives them through the native runtime.
read and reader: inspect observations by role
reader and read give a read-only, role-addressed view of a raw observation. They reuse the model adapter pipeline pointed at the consumer (resolve_from_contract() plus the obs transform with a no-op action), so they are encoding-agnostic across envs and never mutate the observation.
sess.reader(*items) resolves once and returns a callable mapping a raw observation to {role: value}:
import rlmesh.adapters as adapt
with model.session(env) as sess:
read = sess.reader(adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc"), adapt.EEF_POS)
obs, _ = sess.reset(seed=0)
while not sess.done:
view = read(obs) # {IMAGE_PRIMARY: ..., EEF_POS: ...}
screen.show(view[adapt.IMAGE_PRIMARY])
obs, *_ = sess.step(sess.predict(obs))
sess.read(obs, item) is the one-shot single-role convenience. The underlying reader is cached per item, so calling it every step does not re-resolve:
ee = sess.read(obs, adapt.EEF_POS)
img = sess.read(obs, adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc"))
An item is one of:
- A bare role constant (
adapt.EEF_POS,adapt.IMAGE_PRIMARY), kept in the env’s native encoding and using the env’s own declared layout. - A model-input leaf that declares the encoding you want, such as
adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc")oradapt.State(adapt.EEF_POS). The adapter converts to that form whatever the env stores.
Roles and leaves are the same vocabulary the rest of the adapter system uses; see Adapters. The env must publish adapter tags (via an EnvFactory or rlmesh.adapters.tag(...)), or the read raises an AdapterResolutionError, since there are no roles to address otherwise. Values come back in the env’s own framework (NumPy for a Gymnasium env, torch for a torch route).
Reach for it to debug an env: confirm what a camera returns (shape, layout, value range) without threading it through a model. It also gives a consistent way to log canonical roles, recording EEF_POS or the primary image the same way across heterogeneous envs, since the role addresses the quantity rather than the env’s key. The same read works for reward shaping: compute a shaped term over canonical roles, e.g. reward - 0.1 * distance(sess.read(obs, adapt.EEF_POS), goal).
For action chunks and batching, see Performance and Scaling.