Running Evaluations

Run full episodes or drive a model and environment step by step.

Use run() to score a model over complete episodes. Use session() when you need to inspect or control each step. Both drive the same reset, predict, step loop and apply the adapter resolved from the environment’s contract and the model’s ModelSpec.

Pick an entry point

Entry point You get Reach for it when
run() One call drives whole episodes and returns a typed RunResult. Scoring a model: leaderboards, sweeps, CI checks.
session() A Session you step by hand (reset / predict / step). Rendering, custom stop conditions, branching, or mixing your own per-step logic.
Session.reader / Session.read A read-only, role-addressed view of each raw observation. Inspecting an env, logging canonical roles, or shaping a reward.

read and reader let you inspect an observation by role while using a session.

env contractpublished tags
model spec
resolve adapter
run()whole episodes
session()step by hand
read() / reader()inspect by role

run(): the automated rollout

run() completes the episodes and returns a RunResult:

result = model.run(env, seeds=range(100))
print(f"mean reward {result.mean_reward:.2f}")
if result.success_rate is not None:
    print(f"{result.success_rate:.0%} success")

For code you will reuse or package, define an EnvFactory and a Model subclass, then call MyPolicy().run(MyEnv(), episodes=10). env also accepts an existing Gymnasium-style environment, a remote handle, or an address:

result = model.run("tcp://127.0.0.1:5555", seeds=range(100))

The module-level run() accepts a Model class or instance, a remote model, or rlmesh.RANDOM_SAMPLE for an action-space sampling baseline. Wrap a prediction function in the backend Model that matches its array type:

import rlmesh
import rlmesh.numpy

result = rlmesh.run(rlmesh.numpy.Model(my_policy_fn), env, seeds=range(10))
baseline = rlmesh.run(rlmesh.RANDOM_SAMPLE, env, episodes=10)

Arguments

Argument Default Meaning
seeds None One seed per episode; alone, its length sets the episode count.
episodes None Exact episode count; must equal the length of seeds when both are given. 0 runs none.
execution_horizon 1 Actions executed per predicted chunk; see Performance and Scaling.
close_env False Shut the env down when the run finishes (opt-in).
trial_index_base 0 First trial ordinal for environments that declare benchmark trials.

With neither seeds nor episodes, run() does a single episode. execution_horizon is accepted by both the bound methods (model.run / model.session) and the module-level run() / session(), which forwards it through.

model.run() drives the native runtime loop. It accepts hooks= for step and episode callbacks and instruction= for a model with a declared text input. For live viewing, use session() with view=.

Watching and capping the loop

run() also takes max_episode_steps and max_episode_seconds, per-episode caps that mark a capped episode truncated like an env time limit. These caps need the runtime to own resets; use episodes to bound an autoresetting vector env.

Pass hooks= to model.run() or Session.run. A RunHooks subclass can observe on_run_start, on_episode_start, on_step, on_episode_end, and on_run_end. StepEvent includes the step’s observation, action, reward, terminal flags, timings, and a lazy role read. The callbacks run in the same per-episode order on both paths. A hook exception aborts the run, and on_run_end fires once with the completed episodes:

class Progress(rlmesh.RunHooks):
    def on_episode_end(self, result):
        print(f"episode {result.index}: reward {result.reward:.2f}")

result = model.run(env, seeds=range(50), max_episode_steps=500, hooks=Progress())

The result

RunResult is immutable and aggregates its episodes:

Member Type Meaning
.episodes tuple[EpisodeResult, ...] One EpisodeResult per episode.
.mean_reward float Mean total reward across episodes.
.success_rate float | None Fraction the env reported as successful; None if any outcome is unknown.
.num_episodes int Episode count.
.total_steps int Summed steps across episodes.

Each EpisodeResult carries index, seed, steps, reward, terminated, truncated, success, and timing fields (duration_s, predict_ms, step_ms):

for ep in result.episodes:
    print(ep.index, ep.seed, ep.steps, ep.reward, ep.terminated, ep.success)

On the native run path, RunResult.telemetry contains aggregated timing and size measurements. result.format_telemetry() prints them as a table. A Session.run() result has no runtime telemetry rows; its episode timing fields are still available.

session(): manual, step-by-step control

session() hands back a Session you drive yourself. Use it as a context manager so the env connection (and any managed model) closes on exit:

with model.session(env, instruction="put the cup on the plate") as sess:
    obs, info = sess.reset(seed=0)
    while not sess.done:
        action = sess.predict(obs)
        obs, reward, terminated, truncated, info = sess.step(action)

The loop primitives mirror Gymnasium, with the adapter folded in.

  1. reset(seed)
  2. predict(obs)
  3. step(action)
  • not done: back to predict(obs)
  • done: back to reset(seed)
  • sess.reset(seed=None, trial_index=None) → (obs, info). Begins an episode; ends the previous one (firing on_episode_end) and clears adapter state such as the frame-stack buffer. trial_index is the 0-based ordinal of this episode in a benchmark’s trial sweep; it reaches the env as reset(options={"trial_index": ...}), but only if the env declared the key in EnvFactory.reset_options – passing one to an env that did not warns and resets without it. sess.run(trial_index_base=0) walks the ordinals for you (episode i is trial trial_index_base + i), delivers each to a declaring env, and reports each on EpisodeResult.trial whether or not the env asked for it.
  • sess.predict(obs) → action. Applies the model’s adapter around the model’s own predict: the declarative obs transform, host-side frame stacking, any Custom code, instruction injection into declared text leaves, and chunk replay (one action per call). Returns an env-ready action.
  • sess.step(action) → (obs, reward, terminated, truncated, info). Applies the action and records reward and termination.
  • sess.done is True once the current episode terminated or truncated.
  • sess.close() releases the connection, shuts the env down only on the close_env opt-in, and fires on_close.

Use session() when you need to branch on each step, render a frame, or stop early. The module-level session() accepts the same model forms as run(), including rlmesh.RANDOM_SAMPLE.

on_episode_end fires at every episode boundary (the next reset(), or close() for the last episode), so a stateful model clears its per-episode state on either path. Session.run pumps whole episodes through the session primitives; model.run() drives them through the native runtime.

read and reader: inspect observations by role

reader and read give a read-only, role-addressed view of a raw observation. They reuse the model adapter pipeline pointed at the consumer (resolve_from_contract() plus the obs transform with a no-op action), so they are encoding-agnostic across envs and never mutate the observation.

sess.reader(*items) resolves once and returns a callable mapping a raw observation to {role: value}:

import rlmesh.adapters as adapt

with model.session(env) as sess:
    read = sess.reader(adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc"), adapt.EEF_POS)
    obs, _ = sess.reset(seed=0)
    while not sess.done:
        view = read(obs)              # {IMAGE_PRIMARY: ..., EEF_POS: ...}
        screen.show(view[adapt.IMAGE_PRIMARY])
        obs, *_ = sess.step(sess.predict(obs))

sess.read(obs, item) is the one-shot single-role convenience. The underlying reader is cached per item, so calling it every step does not re-resolve:

ee = sess.read(obs, adapt.EEF_POS)
img = sess.read(obs, adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc"))

An item is one of:

  • A bare role constant (adapt.EEF_POS, adapt.IMAGE_PRIMARY), kept in the env’s native encoding and using the env’s own declared layout.
  • A model-input leaf that declares the encoding you want, such as adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc") or adapt.State(adapt.EEF_POS). The adapter converts to that form whatever the env stores.

Roles and leaves are the same vocabulary the rest of the adapter system uses; see Adapters. The env must publish adapter tags (via an EnvFactory or rlmesh.adapters.tag(...)), or the read raises an AdapterResolutionError, since there are no roles to address otherwise. Values come back in the env’s own framework (NumPy for a Gymnasium env, torch for a torch route).

Reach for it to debug an env: confirm what a camera returns (shape, layout, value range) without threading it through a model. It also gives a consistent way to log canonical roles, recording EEF_POS or the primary image the same way across heterogeneous envs, since the role addresses the quantity rather than the env’s key. The same read works for reward shaping: compute a shaped term over canonical roles, e.g. reward - 0.1 * distance(sess.read(obs, adapt.EEF_POS), goal).

For action chunks and batching, see Performance and Scaling.