rlmesh.StepEvent

classfrom rlmesh import StepEvent[source]

One env step of a run, passed to RunHooks.on_step().

StepEvent(
    episode: int,
    seed: int | None,
    step: int,
    observation: Any,
    action: Any,
    reward: float,
    terminated: bool,
    truncated: bool,
    info: Mapping[str, Any],
    predict_ms: float,
    step_ms: float,
    read: Callable[[object], object],
)

The same event on both loops, with one timing definition: step_ms is the wall time of this step’s env round trip, and predict_ms the wall time of the model predict(s) that produced an action since the previous step – the session loop’s synchronous predict; on the native loop the runtime’s own measurement, which is zero for a step served from chunk replay (execution_horizon > 1: the chunk’s predict lands on the step that first consumed it) and, with prefetch_lead > 0, the next chunk’s predict on the step it landed on. On a vectorized env every lane of a group shares both numbers. The per-episode means of these values are EpisodeResult.predict_ms / EpisodeResult.step_ms. A NEXT_STEP autoreset roll is not a step of any episode and never surfaces here.

episode

attribute[source]
episode: int

0-based episode index (equals EpisodeResult.index).

seed

attribute[source]
seed: int | None

The episode’s reset seed, or None.

step

attribute[source]
step: int

0-based step index within the episode.

observation

attribute[source]
observation: Any

The raw observation the action was predicted from (pre-step).

action

attribute[source]
action: Any

The env-ready action applied to the env.

reward

attribute[source]
reward: float

The step’s reward.

terminated

attribute[source]
terminated: bool

Whether the env reported a terminal state on this step.

truncated

attribute[source]
truncated: bool

Whether this step truncated the episode.

info

attribute[source]
info: Mapping[str, Any]

The step’s info mapping.

predict_ms

attribute[source]
predict_ms: float

Wall time of the predict(s) charged to this step, in milliseconds (see above); zero on a chunk-replay step.

step_ms

attribute[source]
step_ms: float

Wall time of the env step round trip, in milliseconds.

read

attribute[source]
read: Callable[[object], object]

Lazy role reader bound to observation – event.read(item) delegates to Session.read(), so resolution is cached per item and never triggered unless called.