rlmesh.StepEvent
One env step of a run, passed to RunHooks.on_step().
StepEvent(
episode: int,
seed: int | None,
step: int,
observation: Any,
action: Any,
reward: float,
terminated: bool,
truncated: bool,
info: Mapping[str, Any],
predict_ms: float,
step_ms: float,
read: Callable[[object], object],
)The same event on both loops, with one timing definition: step_ms is
the wall time of this step’s env round trip, and predict_ms the wall
time of the model predict(s) that produced an action since the previous
step – the session loop’s synchronous predict; on the native loop the
runtime’s own measurement, which is zero for a step served from chunk
replay (execution_horizon > 1: the chunk’s predict lands on the step
that first consumed it) and, with prefetch_lead > 0, the next chunk’s
predict on the step it landed on. On a vectorized env every lane of a
group shares both numbers. The per-episode means of these values are
EpisodeResult.predict_ms / EpisodeResult.step_ms. A
NEXT_STEP autoreset roll is not a step of any episode and never
surfaces here.
Attributes
- episode0-based episode index (equals
EpisodeResult.index). - seedThe episode's reset seed, or
None. - step0-based step index within the episode.
- observationThe raw observation the action was predicted from (pre-step).
- actionThe env-ready action applied to the env.
- rewardThe step's reward.
- terminatedWhether the env reported a terminal state on this step.
- truncatedWhether this step truncated the episode.
- infoThe step's
infomapping. - predict_msWall time of the predict(s) charged to this step, in milliseconds (see above); zero on a chunk-replay step.
- step_msWall time of the env
stepround trip, in milliseconds. - readLazy role reader bound to
observation--event.read(item)delegates toSession.read(), so resolution is cached per item and never triggered unless called.
episode
attribute[source]episode: int0-based episode index (equals EpisodeResult.index).
seed
attribute[source]seed: int | NoneThe episode’s reset seed, or None.
step
attribute[source]step: int0-based step index within the episode.
observation
attribute[source]observation: AnyThe raw observation the action was predicted from (pre-step).
action
attribute[source]action: AnyThe env-ready action applied to the env.
reward
attribute[source]reward: floatThe step’s reward.
terminated
attribute[source]terminated: boolWhether the env reported a terminal state on this step.
truncated
attribute[source]truncated: boolWhether this step truncated the episode.
info
attribute[source]info: Mapping[str, Any]The step’s info mapping.
predict_ms
attribute[source]predict_ms: floatWall time of the predict(s) charged to this step, in milliseconds (see above); zero on a chunk-replay step.
step_ms
attribute[source]step_ms: floatWall time of the env step round trip, in milliseconds.
read
attribute[source]read: Callable[[object], object]Lazy role reader bound to observation – event.read(item) delegates to Session.read(), so resolution is cached per item and never triggered unless called.