rlmesh.EpisodeResult
The outcome of one evaluation episode.
EpisodeResult(
index: int,
seed: int | None,
steps: int,
reward: float,
terminated: bool,
truncated: bool,
success: bool | None = None,
duration_s: float = 0.0,
predict_ms: float | None = None,
step_ms: float | None = None,
trial: int | None = None,
)Attributes
- index0-based episode index within the run.
- seedThe seed the episode was reset with, or
None. - stepsNumber of env steps taken.
- rewardTotal reward accumulated over the episode.
- terminatedWhether the env reported a terminal state.
- truncatedWhether the episode was cut short (env truncation, a
max_episode_steps/max_episode_secondscap, or the built-in step bound). - successThe env-reported task outcome from the final step's
info(Gymnasium'sis_success/successkey), orNonewhen the env emits no such signal. - duration_sWall time from reset-return to episode end, in seconds.
- predict_msThis episode's mean per-step wall time of
predict, in milliseconds -- the mean of itsStepEvent.predict_msvalues, so on the native loop a chunk's predict is spread over the steps it served. - step_msThis episode's mean per-step wall time of the env
stepround trip, in milliseconds, orNoneon the same terms. - trialThe trial ordinal the episode walked,
trial_index_base + indexonrun().
index
attribute[source]index: int0-based episode index within the run.
seed
attribute[source]seed: int | NoneThe seed the episode was reset with, or None.
steps
attribute[source]steps: intNumber of env steps taken.
reward
attribute[source]reward: floatTotal reward accumulated over the episode.
terminated
attribute[source]terminated: boolWhether the env reported a terminal state.
truncated
attribute[source]truncated: boolWhether the episode was cut short (env truncation, a max_episode_steps / max_episode_seconds cap, or the built-in step bound).
success
attribute[source]success: bool | NoneThe env-reported task outcome from the final step’s info (Gymnasium’s is_success / success key), or None when the env emits no such signal. Distinct from terminated (which only says the episode reached a terminal state, not how it ended); an unknown outcome is never inferred from it.
duration_s
attribute[source]duration_s: floatWall time from reset-return to episode end, in seconds.
predict_ms
attribute[source]predict_ms: float | NoneThis episode’s mean per-step wall time of predict, in milliseconds – the mean of its StepEvent.predict_ms values, so on the native loop a chunk’s predict is spread over the steps it served. None when no step was timed (an episode that ended before its first step); never a run-wide average.
step_ms
attribute[source]step_ms: float | NoneThis episode’s mean per-step wall time of the env step round trip, in milliseconds, or None on the same terms.
trial
attribute[source]trial: int | NoneThe trial ordinal the episode walked, trial_index_base + index on run(). Delivered as reset(options={"trial_index": ...}) only to an env that declared the option, but recorded either way so the sweep can be read off the result. None for a hand-driven Session.reset() that passed no trial_index, and for an env that owns its resets (NEXT_STEP autoreset on the native loop).