rlmesh.EpisodeResult

classfrom rlmesh import EpisodeResult[source]

The outcome of one evaluation episode.

EpisodeResult(
    index: int,
    seed: int | None,
    steps: int,
    reward: float,
    terminated: bool,
    truncated: bool,
    success: bool | None = None,
    duration_s: float = 0.0,
    predict_ms: float | None = None,
    step_ms: float | None = None,
    trial: int | None = None,
)

index

attribute[source]
index: int

0-based episode index within the run.

seed

attribute[source]
seed: int | None

The seed the episode was reset with, or None.

steps

attribute[source]
steps: int

Number of env steps taken.

reward

attribute[source]
reward: float

Total reward accumulated over the episode.

terminated

attribute[source]
terminated: bool

Whether the env reported a terminal state.

truncated

attribute[source]
truncated: bool

Whether the episode was cut short (env truncation, a max_episode_steps / max_episode_seconds cap, or the built-in step bound).

success

attribute[source]
success: bool | None

The env-reported task outcome from the final step’s info (Gymnasium’s is_success / success key), or None when the env emits no such signal. Distinct from terminated (which only says the episode reached a terminal state, not how it ended); an unknown outcome is never inferred from it.

duration_s

attribute[source]
duration_s: float

Wall time from reset-return to episode end, in seconds.

predict_ms

attribute[source]
predict_ms: float | None

This episode’s mean per-step wall time of predict, in milliseconds – the mean of its StepEvent.predict_ms values, so on the native loop a chunk’s predict is spread over the steps it served. None when no step was timed (an episode that ended before its first step); never a run-wide average.

step_ms

attribute[source]
step_ms: float | None

This episode’s mean per-step wall time of the env step round trip, in milliseconds, or None on the same terms.

trial

attribute[source]
trial: int | None

The trial ordinal the episode walked, trial_index_base + index on run(). Delivered as reset(options={"trial_index": ...}) only to an env that declared the option, but recorded either way so the sweep can be read off the result. None for a hand-driven Session.reset() that passed no trial_index, and for an env that owns its resets (NEXT_STEP autoreset on the native loop).