Adapters

Match independently authored models and environments by declared roles.

rlmesh.adapters derives a model-to-environment IO adapter at runtime from two declarations: an environment tags its observation and action spaces, a model specifies the payload it ingests, and resolve() matches them by role. This replaces most of the per-(model, environment) glue you would otherwise write by hand. Cases the declarative specs do not cover fall back to an escape hatch (see Escape Hatches).

It is opt-in. Nothing here is imported by the core Gymnasium loop. Direct adapter calls and the examples below use the NumPy backend (pip install "rlmesh[numpy]"); model runtime paths use the active RLMesh backend.

Use the Adapter Reference for roles, fields, and conversion rules. See Escape Hatches when a pairing needs custom code.

The core idea

The two sides of an eval are declared independently and never import each other. An environment publishes tags, a model declares a spec, and resolve bridges them by matching semantic roles.

Env tagsroles + a few factsmatches by role
obs / action spaceswidths / dtypes
Model specfull payload + action layoutmatches by role
resolve()→ Adapter
transform_obs
transform_action

The asymmetry between the two sides is deliberate.

Side What it declares
Environment (EnvTags) Each space entry’s role, plus the few facts a gymnasium space cannot carry: image axis layout, rotation encoding, an explicit range. Keys, widths, dtypes, and bounds are read from the spaces.
Model (ModelSpec) The full payload it ingests and the action it emits, in its own conventions: sizes, encodings, container shapes.

Roles match exactly between the environment and model. Use the built-in constants such as IMAGE_PRIMARY, EEF_POS, and EEF_ROT for shared conventions, or an x/ role for a custom pair. Role prefixes and validation rules are covered in the Adapter Reference. Widths, dtypes, and bounds come from the environment’s spaces.

Tag the environment

An environment tags its observation and action spaces once. The role is the first argument on every tag; everything else is the facts the spaces cannot carry.

import rlmesh.adapters as adapt

tags = adapt.EnvTags(
    observation={
        "wrist_rgb": adapt.ImageTag(adapt.IMAGE_PRIMARY),
        "ee_pos": adapt.StateTag(adapt.EEF_POS),
        "ee_quat": adapt.StateTag(adapt.EEF_ROT, encoding="quat_xyzw"),
        "grip": adapt.StateTag(adapt.GRIPPER_POS),
        "goal": adapt.TextTag(adapt.INSTRUCTION),
    },
    action=adapt.Action(
        adapt.Actuator(adapt.ACTION_DELTA_POS, dim=3),
        adapt.Actuator(adapt.ACTION_DELTA_ROT, dim=3, encoding="axis_angle"),
        adapt.Actuator(adapt.ACTION_GRIPPER, dim=1, range=(-1.0, 1.0)),
        clip=(-1.0, 1.0),
    ),
)

The observation is a tree whose container is the runtime container: a dict maps a Dict space, a tuple maps a Tuple space, and a bare leaf tags a single space leaf. Nesting is real dict nesting that mirrors a nested Dict space ({"agent": {"eef_pos": adapt.StateTag(adapt.EEF_POS)}}), not dotted keys.

When you author an environment with Environments, the EnvFactory tags class attribute is stamped onto the env automatically, so the same tags ride a local env and a served one.

Flat (non-Dict) observations

Some environments expose a single flat numeric vector with fixed index ranges instead of one key per quantity (Metaworld is the common case). A Split tags that vector. It is the observation-side mirror of Action: a sequence of Field slices in order, each naming its role with offsets implied by order. A field with no role is a skip that advances the offset over indices the model does not read.

"proprio": adapt.Split(
    adapt.Field(adapt.EEF_POS, 3),
    adapt.Field(adapt.EEF_ROT, 4, encoding="quat_xyzw"),
    adapt.Field(adapt.GRIPPER_POS, 1),
    adapt.Field(dim=10),  # object/goal indices the policy reads from pixels
),

Split is a leaf, not a container. When the whole observation is one flat box, pass it directly:

adapt.EnvTags(observation=adapt.Split(...), action=adapt.Action(...))

A model matches purely by role, so the same spec resolves against a flat env and a Dict env with no change. The full Field table is in Adapter Reference.

Specify the model

A model fully specifies the payload it ingests and the action it emits, in its own conventions. The role is again the first argument; size= sets a square image’s height and width together.

spec = adapt.ModelSpec(
    input={
        "image": adapt.Image(adapt.IMAGE_PRIMARY, size=224),
        "proprio": adapt.Concat(
            adapt.EEF_POS,
            adapt.State(adapt.EEF_ROT, encoding="rot6d"),
            adapt.GRIPPER_POS,
        ),
        "task": adapt.Text(adapt.INSTRUCTION),
    },
    output=adapt.Action(
        adapt.Actuator(adapt.ACTION_DELTA_POS, dim=3),
        adapt.Actuator(adapt.ACTION_DELTA_ROT, dim=6, encoding="rot6d"),
        adapt.Actuator(adapt.ACTION_GRIPPER, dim=1, range=(-1.0, 1.0)),
    ),
)

The input is a tree whose container is the payload the prediction function receives: a dict (each key a payload slot), a tuple, or a bare single leaf. A leaf carries no key: its position in the tree is the payload position, and a role may be reused across leaves. Concat is the multi-part state leaf: each part is a bare role string (sugar for a role-only State) or a State, concatenated in order. Every leaf and its options are enumerated in Adapter Reference.

Resolve and apply

resolve() matches the model spec against the tags and the spaces and returns an Adapter. The adapter preprocesses an observation into the model’s input format and postprocesses the model’s action back into the environment’s.

adapter = adapt.resolve(tags, env.observation_space, env.action_space, spec)
print(adapter.explain())               # the exact transforms chosen
payload = adapter.transform_obs(obs)    # env observation -> model input
action = adapter.transform_action(out)  # model output    -> env action

explain() prints what the resolver derived. For the pair above the image is resized, the rotation goes quat_xyzw -> rot6d, the instruction key is remapped (goal -> task), and the 6-d rotation in the model’s action is converted back to the env’s 3-d axis_angle and clipped. Resolution raises AdapterResolutionError when a model input or action actuator has no usable counterpart, or when a declared conversion is impossible. The conversion policy in Adapter Reference decides which conversions apply silently, warn, or fail.

Run a model with no glue

The shortest path publishes the tags on the served environment and lets the model resolve the adapter from the contract.

server = rlmesh.EnvServer(env, "127.0.0.1:5555", tags=tags)
server.serve()

EnvServer(tags=...) validates the tags against the environment’s spaces and merges them into the contract metadata (the tag() verb does the same for an environment you serve yourself). A model then resolves from the handshake alone.

from rlmesh.numpy import Model, RemoteEnv

env = RemoteEnv("127.0.0.1:5555")
model = Model(predict, spec=spec)  # predict works in the model's own format
model.run(env, episodes=10)

run(env) reads the environment’s contract, resolves the adapter, and wraps predict so it only ever sees the model’s declared payload. To resolve explicitly, use resolve_from_contract() and adapter.wrap_predict(predict). See Models for the prediction corners a predict may implement, and Serve an Environment for addresses, readiness, and health. examples/python/adapters is the smallest end-to-end serve-and-run loop.

Frame history

A model that conditions on a short history of frames declares stack=N on an image input. The adapter buffers the last N processed frames host-side and emits them on a new leading axis ((N, H, W, C)), padding the start of an episode with the first frame and clearing the buffer on reset.

"image": adapt.Image(adapt.IMAGE_PRIMARY, size=224, stack=4)

At execution_horizon=1, frame stacking adds no observation traffic. At higher horizons, a served model also receives the observations from replayed steps so its history stays complete. See Performance and Scaling.

Known limitations

The system targets the manipulation/VLA case: RGB cameras, proprioception, and an instruction. A few things are out of scope for now and fall back to an escape hatch.

Area Status
Modalities beyond image / state / text Depth, lidar, and point clouds are not first-class; carry them through a Custom input or a custom AdapterBase.
Tokenization Stays in the model. Text delivers the instruction as a string; tokenize it inside your prediction function. There is intentionally no TokenizerInput.
Rotation encodings Fixed set: quat_xyzw, quat_wxyz, axis_angle, rot6d, rot6d_rowmajor, euler_xyz. For a one-off convention, declare a CustomEncoding; see Escape Hatches.