rlmesh.adapters.AdapterBase
Base class for env-to-model adapters.
rlmesh.adapters.resolve() derives the spec-driven implementation
(Adapter); subclass this directly to plug a fully custom
pairing into anything built around adapters instead. Implement the two
transforms; wrap_predict comes for free. Custom adapters may hold
state across steps (e.g. cache proprio from transform_obs for use
in transform_action) – that is their power over declarative specs.
Override reset() to clear any such state at episode boundaries.
Methods
- transform_obs()Convert a raw env observation into the model input payload.
- transform_action()Convert a model action output into the env action.
- reset()Clear episode-scoped state at an episode boundary.
- observe()Advance episode-scoped state for a step that predicted nothing.
- explain()Return a human-readable summary of the adapter.
- wrap_predict()Wrap a model predict function with both transforms.
transform_obs
function[source]transform_obs(raw_obs: RawObs) -> AnyConvert a raw env observation into the model input payload.
raw_obs is a mapping for a Dict space, or a bare array/leaf for a
flat (non-Dict) space. The return is the model input payload tree (a
dict, a list, or a bare leaf – whatever the model spec’s input tree
declares).
transform_action
function[source]transform_action(raw_action: object) -> ActionTConvert a model action output into the env action.
reset
function[source]reset() -> NoneClear episode-scoped state at an episode boundary.
The base adapter holds no episode-scoped state; a stateful custom adapter overrides this to clear its per-episode state. It is driven on the single-env local loop, so there is no per-lane bookkeeping to do (the served path stacks natively with episode-keyed buffers in the core).
observe
function[source]observe(raw_obs: RawObs) -> NoneAdvance episode-scoped state for a step that predicted nothing.
Driven once per env step whose action came from a replayed chunk: the
step happened, so an adapter holding a frame history still has to see
its observation or the window would hold decision points instead of
consecutive steps. A no-op by default – override it alongside
reset() if your adapter caches anything across steps.
explain
function[source]explain() -> strReturn a human-readable summary of the adapter.
wrap_predict
function[source]wrap_predict(
predict_fn: Callable[[Any], object],
) -> Callable[[Any], ActionT]Wrap a model predict function with both transforms.
predict_fn receives the model input payload tree – a dict for a
Dict-shaped input, a list for a Tuple-shaped input, or a bare leaf for a
single-leaf input (whatever transform_obs() produces). The
returned callable takes a raw env observation – a mapping, or a
bare array/leaf for a flat (non-Dict) env – and returns an env-ready
action, suitable for rlmesh.numpy.Model.
The execution horizon is a runtime decision (execution_horizon on
ResolveAdapter, owned by the runtime driver) and the served engine
emits the chunk. This direct wrapper applies one action per step;
in-process chunk replay lives in
rlmesh._models._chunk.ChunkReplay, driven by run with a
locally chosen horizon.