Escape Hatches

Add custom input, rotation encoding, or adapter behavior when a spec is not enough.

Start with a declarative adapter spec. When one part of a model and environment pairing needs Python code, choose the smallest extension that covers it. Custom changes one input; a pair override replaces the whole adapter.

Hatch Scope Reach for it when
Custom input one payload slot one input is computed from the raw observation; the rest stays declarative
CustomEncoding one rotation field a model uses a rotation packing RLMesh does not ship as a built-in
AdapterBase subclass the whole adapter, model-wide the model needs state across steps (e.g. temporal ensembling)
Pair override one (model, env) pairing a single pairing is special enough to hand-write end to end

Custom input: one payload slot

A Custom leaf fills one slot of the model payload from the raw observation with host code. Everything else in the input tree stays spec-driven. Use it for a modality RLMesh does not model first-class (depth, lidar, point clouds) or any value that needs bespoke assembly.

The leaf takes a callable that receives the raw observation mapping and returns the slot’s value:

import rlmesh.adapters as adapt

def encode_depth(obs):
    return obs["depth"].astype("float32") / 1000.0

spec = adapt.ModelSpec(
    input={
        "image": adapt.Image(adapt.IMAGE_PRIMARY, size=224),
        "proprio": adapt.Concat(adapt.EEF_POS, adapt.GRIPPER_POS),
        "depth": adapt.Custom(transform=encode_depth),
    },
    output=adapt.Action(
        adapt.Actuator(adapt.ACTION_DELTA_POS, dim=3),
        adapt.Actuator(adapt.ACTION_GRIPPER, dim=1),
    ),
)

transform= is local only: a callable cannot be serialized, so a spec carrying one cannot be published in contract metadata and resolves only in the process that defined it. To load a transform by name in the local process, use an entrypoint:

adapt.Custom(entrypoint="my_pkg.adapters:encode_depth")

Neither form can be published in v1 contract metadata. The entrypoint is imported only under resolve() with trust_entrypoints=True; otherwise resolution refuses to import it. Set exactly one of transform= or entrypoint=.

Custom encoding: a host-side rotation packing

Rotation encodings are a closed vocabulary so a spec can resolve on a remote client with no code. When a model expects a packing outside that set, a CustomEncoding layers a host-side repack on top of the nearest built-in base encoding, keeping the rest of the rotation pipeline declarative.

You supply the two arms of an encode/decode pair and drop the result in as an encoding=:

import rlmesh.adapters as adapt

rot6d_packed = adapt.CustomEncoding(
    "rot6d",
    from_base=unpack_rot6d,  # base -> custom, applied on the obs side
    to_base=pack_rot6d,      # custom -> base, applied on the action side
    name="rot6d_packed",
)

spec = adapt.ModelSpec(
    input={"proprio": adapt.Concat(adapt.EEF_POS, adapt.State(adapt.EEF_ROT, encoding=rot6d_packed))},
    output=adapt.Action(adapt.Actuator(adapt.ACTION_DELTA_ROT, dim=6, encoding=rot6d_packed)),
)

The field resolves natively as its base: role matching, range mapping, and the env-to-base rotation conversion are unchanged. The adapter then repacks at the field boundary: from_base (base -> custom) on the observation side and to_base (custom -> base) on the action side, so the model sees its packing and the env sees the base.

Rule Detail
Width-preserving the packing must keep the base width (ROTATION_DIMS[base]); it repacks, it does not resize
Arm per side from_base is needed only when the encoding tags an observation state; to_base only when it tags an action; supply at least one
Arms agree both in-process callables, or both "module:callable" entrypoint strings; never a mix
Schema, not runtime a CustomEncoding serializes as a describe/validate schema ({base, from_base, to_base}): the control plane and dashboard show it and resolve its base against an env, but never run the arm. Entrypoint arms travel as their module:callable string; an in-process callable has no wire form, so it travels as a non-portable <local> marker
Any offset the part it tags may sit anywhere in a multi-part Concat: the repack reads and writes exactly its own slice, whose offset comes from the resolved plan (part widths are env-dependent, so they are known only after resolve). A 20-wide bimanual proprio can carry one per arm
Executes once the transform runs in exactly one place — the process that defined the in-process callable. Running it from a serialized spec, or from an entrypoint reference, is a hard failure: the arm is language-tied and is not relocated or injected. Keep the callable in the process that executes the transform

Use a CustomEncoding for a one-off or proprietary packing you apply where the model lives. Its schema travels so the platform can show the ModelSpec and verify base-vs-env, while the transform stays pinned to its defining process. When the convention is instead general, stable, and published (a checkpoint’s documented rotation format), upstream the packing to the first-party RotationEncoding enum: it then resolves by role with no host code at all, is conformance-tested, and — being fully native — the whole adapter can relocate off the model layer (e.g. into middleman).

Wrong-looking rotations: which hatch

A rotation that comes out wrong is one of two things, and they have different answers:

  • A rigid, data-independent re-orientation — a wrist camera or gripper mounted turned, an embodiment whose tool frame is the env’s rotated by a constant. That is a fixed rotation composed onto every value, so it belongs in the spec as data, not as code: post_rotate on the State (a Rotation literal, right-multiplied onto the resolved rotation). Build it with Rotation.from_matrix(rows); see Adapter Reference.
  • A data-dependent repack — the model’s convention packs the same rotation differently, or a checkpoint was trained against a producer that read its quaternion in another component order. No fixed rotation describes that, so it is a CustomEncoding: the field still resolves as its base, and the two arms repack at the field boundary.

Not a hatch: latching an action across steps

A sticky gripper holds the last commanded value for a fixed number of environment steps. Implement that in the environment wrapper. The adapter transforms an entire action chunk when the model predicts; it does not run again for each replayed step. Counting predictions in an adapter will therefore give the wrong timing when chunks contain multiple actions.

AdapterBase subclass: stateful behavior

For state beyond the built-in frame and previous-action history, such as ACT-style temporal ensembling, subclass AdapterBase and wrap a resolved adapter, adding only the stateful part. Observation handling and per-step action conversion (encodings, ranges, clipping) stay spec-driven; the custom code is the state alone.

import rlmesh.adapters as adapt
from collections import deque

class ChunkEnsembleAdapter(adapt.AdapterBase):
    def __init__(self, inner: adapt.Adapter, horizon: int = 8):
        self._inner = inner
        self._chunks: deque = deque(maxlen=horizon)

    def transform_obs(self, raw_obs):
        return self._inner.transform_obs(raw_obs)

    def transform_action(self, raw_action):
        ...  # remember the chunk, ensemble every live chunk's row for this step
        return self._inner.transform_action(ensembled)  # spec-driven conversion

    def reset(self):
        self._chunks.clear()  # never ensemble across episodes

Build it by resolving first, then wrapping:

adapter = ChunkEnsembleAdapter(adapt.resolve(tags, env.observation_space, env.action_space, spec))

A stateful custom adapter keeps affinity to its stream: override reset to clear episode-scoped state. It is wired to the per-episode boundary (the model’s on_episode_end) so a finished episode never bleeds into the next, and it runs on the single-env local loop, so there is no per-lane bookkeeping to do. On the served path, frame-stacking is handled natively in the core with episode-keyed buffers, so a vectorized route needs no per-adapter affinity.

The full worked example (the spec, the ensemble math, and the factory that resolves then wraps) is examples/python/vla_adapters/models/act.py.

Pair override: replace the adapter for one pairing

When a single (model, env) pairing is special enough that none of the above fit, replace its adapter outright. This needs no RLMesh machinery: keep a registry keyed by the pair and consult it before resolve().

import rlmesh.adapters as adapt

OVERRIDES = {
    ("xvla", "simpler-bridge"): XVLABridgeAdapter,
}

def build_adapter(model_name, env_name, env, spec):
    if (override := OVERRIDES.get((model_name, env_name))) is not None:
        return override()
    return adapt.resolve(env.tags, env.observation_space, env.action_space, spec)

Because both the resolved adapter and any override are AdapterBase instances, the rest of the eval loop (wrap_predict, the served path via resolve_from_contract()) is identical either way. An override is only for pairing-wide special-casing; a per-slot or per-field need belongs in one of the earlier hatches.