Escape Hatches
Add custom input, rotation encoding, or adapter behavior when a spec is not enough.
Start with a declarative adapter spec. When one part of a model and environment pairing needs Python code, choose the smallest extension that covers it. Custom changes one input; a pair override replaces the whole adapter.
| Hatch | Scope | Reach for it when |
|---|---|---|
Custom input |
one payload slot | one input is computed from the raw observation; the rest stays declarative |
CustomEncoding |
one rotation field | a model uses a rotation packing RLMesh does not ship as a built-in |
AdapterBase subclass |
the whole adapter, model-wide | the model needs state across steps (e.g. temporal ensembling) |
| Pair override | one (model, env) pairing |
a single pairing is special enough to hand-write end to end |
Custom input: one payload slot
A Custom leaf fills one slot of the model payload from the raw observation with host code. Everything else in the input tree stays spec-driven. Use it for a modality RLMesh does not model first-class (depth, lidar, point clouds) or any value that needs bespoke assembly.
The leaf takes a callable that receives the raw observation mapping and returns the slot’s value:
import rlmesh.adapters as adapt
def encode_depth(obs):
return obs["depth"].astype("float32") / 1000.0
spec = adapt.ModelSpec(
input={
"image": adapt.Image(adapt.IMAGE_PRIMARY, size=224),
"proprio": adapt.Concat(adapt.EEF_POS, adapt.GRIPPER_POS),
"depth": adapt.Custom(transform=encode_depth),
},
output=adapt.Action(
adapt.Actuator(adapt.ACTION_DELTA_POS, dim=3),
adapt.Actuator(adapt.ACTION_GRIPPER, dim=1),
),
)
transform= is local only: a callable cannot be serialized, so a spec carrying one cannot be published in contract metadata and resolves only in the process that defined it. To load a transform by name in the local process, use an entrypoint:
adapt.Custom(entrypoint="my_pkg.adapters:encode_depth")
Neither form can be published in v1 contract metadata. The entrypoint is imported only under resolve() with trust_entrypoints=True; otherwise resolution refuses to import it. Set exactly one of transform= or entrypoint=.
Custom encoding: a host-side rotation packing
Rotation encodings are a closed vocabulary so a spec can resolve on a remote client with no code. When a model expects a packing outside that set, a CustomEncoding layers a host-side repack on top of the nearest built-in base encoding, keeping the rest of the rotation pipeline declarative.
You supply the two arms of an encode/decode pair and drop the result in as an encoding=:
import rlmesh.adapters as adapt
rot6d_packed = adapt.CustomEncoding(
"rot6d",
from_base=unpack_rot6d, # base -> custom, applied on the obs side
to_base=pack_rot6d, # custom -> base, applied on the action side
name="rot6d_packed",
)
spec = adapt.ModelSpec(
input={"proprio": adapt.Concat(adapt.EEF_POS, adapt.State(adapt.EEF_ROT, encoding=rot6d_packed))},
output=adapt.Action(adapt.Actuator(adapt.ACTION_DELTA_ROT, dim=6, encoding=rot6d_packed)),
)
The field resolves natively as its base: role matching, range mapping, and the env-to-base rotation conversion are unchanged. The adapter then repacks at the field boundary: from_base (base -> custom) on the observation side and to_base (custom -> base) on the action side, so the model sees its packing and the env sees the base.
| Rule | Detail |
|---|---|
| Width-preserving | the packing must keep the base width (ROTATION_DIMS[base]); it repacks, it does not resize |
| Arm per side | from_base is needed only when the encoding tags an observation state; to_base only when it tags an action; supply at least one |
| Arms agree | both in-process callables, or both "module:callable" entrypoint strings; never a mix |
| Schema, not runtime | a CustomEncoding serializes as a describe/validate schema ({base, from_base, to_base}): the control plane and dashboard show it and resolve its base against an env, but never run the arm. Entrypoint arms travel as their module:callable string; an in-process callable has no wire form, so it travels as a non-portable <local> marker |
| Any offset | the part it tags may sit anywhere in a multi-part Concat: the repack reads and writes exactly its own slice, whose offset comes from the resolved plan (part widths are env-dependent, so they are known only after resolve). A 20-wide bimanual proprio can carry one per arm |
| Executes once | the transform runs in exactly one place — the process that defined the in-process callable. Running it from a serialized spec, or from an entrypoint reference, is a hard failure: the arm is language-tied and is not relocated or injected. Keep the callable in the process that executes the transform |
Use a CustomEncoding for a one-off or proprietary packing you apply where the model lives. Its schema travels so the platform can show the ModelSpec and verify base-vs-env, while the transform stays pinned to its defining process. When the convention is instead general, stable, and published (a checkpoint’s documented rotation format), upstream the packing to the first-party RotationEncoding enum: it then resolves by role with no host code at all, is conformance-tested, and — being fully native — the whole adapter can relocate off the model layer (e.g. into middleman).
Wrong-looking rotations: which hatch
A rotation that comes out wrong is one of two things, and they have different answers:
- A rigid, data-independent re-orientation — a wrist camera or gripper mounted turned, an embodiment whose tool frame is the env’s rotated by a constant. That is a fixed rotation composed onto every value, so it belongs in the spec as data, not as code:
post_rotateon theState(aRotationliteral, right-multiplied onto the resolved rotation). Build it withRotation.from_matrix(rows); see Adapter Reference. - A data-dependent repack — the model’s convention packs the same rotation differently, or a checkpoint was trained against a producer that read its quaternion in another component order. No fixed rotation describes that, so it is a
CustomEncoding: the field still resolves as its base, and the two arms repack at the field boundary.
Not a hatch: latching an action across steps
A sticky gripper holds the last commanded value for a fixed number of environment steps. Implement that in the environment wrapper. The adapter transforms an entire action chunk when the model predicts; it does not run again for each replayed step. Counting predictions in an adapter will therefore give the wrong timing when chunks contain multiple actions.
AdapterBase subclass: stateful behavior
For state beyond the built-in frame and previous-action history, such as ACT-style temporal ensembling, subclass AdapterBase and wrap a resolved adapter, adding only the stateful part. Observation handling and per-step action conversion (encodings, ranges, clipping) stay spec-driven; the custom code is the state alone.
import rlmesh.adapters as adapt
from collections import deque
class ChunkEnsembleAdapter(adapt.AdapterBase):
def __init__(self, inner: adapt.Adapter, horizon: int = 8):
self._inner = inner
self._chunks: deque = deque(maxlen=horizon)
def transform_obs(self, raw_obs):
return self._inner.transform_obs(raw_obs)
def transform_action(self, raw_action):
... # remember the chunk, ensemble every live chunk's row for this step
return self._inner.transform_action(ensembled) # spec-driven conversion
def reset(self):
self._chunks.clear() # never ensemble across episodes
Build it by resolving first, then wrapping:
adapter = ChunkEnsembleAdapter(adapt.resolve(tags, env.observation_space, env.action_space, spec))
A stateful custom adapter keeps affinity to its stream: override reset to clear episode-scoped state. It is wired to the per-episode boundary (the model’s on_episode_end) so a finished episode never bleeds into the next, and it runs on the single-env local loop, so there is no per-lane bookkeeping to do. On the served path, frame-stacking is handled natively in the core with episode-keyed buffers, so a vectorized route needs no per-adapter affinity.
The full worked example (the spec, the ensemble math, and the factory that resolves then wraps) is examples/python/vla_adapters/models/act.py.
Pair override: replace the adapter for one pairing
When a single (model, env) pairing is special enough that none of the above fit, replace its adapter outright. This needs no RLMesh machinery: keep a registry keyed by the pair and consult it before resolve().
import rlmesh.adapters as adapt
OVERRIDES = {
("xvla", "simpler-bridge"): XVLABridgeAdapter,
}
def build_adapter(model_name, env_name, env, spec):
if (override := OVERRIDES.get((model_name, env_name))) is not None:
return override()
return adapt.resolve(env.tags, env.observation_space, env.action_space, spec)
Because both the resolved adapter and any override are AdapterBase instances, the rest of the eval loop (wrap_predict, the served path via resolve_from_contract()) is identical either way. An override is only for pairing-wide special-casing; a per-slot or per-field need belongs in one of the earlier hatches.