Troubleshooting
Diagnose adapter resolution, transport, and runtime failures.
Start by checking whether the failure happens during adapter resolution or after the connection is established. Resolution errors usually point to a mismatch between the model spec and environment tags. Runtime errors can come from a failed connection, an environment, or the model itself.
The sections below explain the exceptions and recovery steps. If a pairing runs but produces unexpected observations or actions, use the adapter inspection tools.
The exception families
Two families cover almost everything.
Resolution failures raise AdapterResolutionError, a subclass of ValueError. It comes from resolve(), resolve_from_contract(), tag(), the read API, and the adapter resolution that run/session do on connect. It always points at the offending leaf and what it expected, and it fires before a single step runs.
Runtime failures come from the native core and reach Python through a small exception hierarchy plus a few standard built-ins. The native module defines RLMeshException (a subclass of RuntimeError) as the base, with ProtocolException and EnvironmentException beneath it. An environment that reports a fault while serving a request raises EnvironmentException. Transport faults, timeouts, and bad arguments map to the standard ConnectionError, TimeoutError, and ValueError instead, so ordinary except clauses catch them without importing anything RLMesh-specific.
What each call raises
resolve() and resolve_from_contract()
Both raise AdapterResolutionError when the model spec cannot be bridged to the env’s tags and spaces: a required role with no optional fallback, a declared channel mismatch, an upscale without allow_upscale, an aspect mismatch without fit, an unsupported resample or dtype, an impossible encoding conversion, an unknown field on a known leaf, or a join-time class/width/encoding/range disagreement. resolve_from_contract() adds two of its own: the contract carries no adapter tags, or those tags are not serializable JSON.
import rlmesh.adapters as adapt
try:
adapter = adapt.resolve(tags, env.observation_space, env.action_space, spec)
except adapt.AdapterResolutionError as exc:
print(exc) # names the leaf and what it expected
A spec that references a Custom input by entrypoint= also raises here unless you pass resolve(..., trust_entrypoints=True). The conversion-policy table in Adapter Reference decides which conversions apply silently, warn, or fail.
run(), session(), and Session.predict()
The eval loop resolves the adapter on the first connect, so a spec/env mismatch raises AdapterResolutionError before any episode begins. A few more checks run at the same point:
- An env that publishes adapter tags paired with a model whose
specisNoneraisesAdapterResolutionError: pass aModelSpec, orrlmesh.NO_ADAPTERif the model adapts its own observations. - A target that is neither an env, an
EnvFactory, a remote handle, nor an address string raisesTypeError.
Once running, predict and step surface whatever the env or model raises. For a local model, predict runs in-process, so a bug in your prediction function propagates as its own native Python exception. For a served model, the server maps a handler that declines a request to RuntimeError ("model error: ...").
serve()
EnvServer validates published tags against the env’s spaces at construction (for a scalar env), so a bad tag raises AdapterResolutionError at startup instead of when the first model connects. Asking for device= on a numpy env, or combining an explicit address with host/port/path, raises ValueError. A bind that fails surfaces as RuntimeError ("server error: ...").
A served model resolves its adapter once per env, at the configure step rather than at connect, so a spec/env mismatch fails route configuration loudly with AdapterResolutionError rather than predicting wrongly.
Symptom, cause, fix: resolution and transport
| Symptom | Cause | Fix |
|---|---|---|
AdapterResolutionError: ... no usable counterpart |
a model input or actuator role the env never tags | tag the role on the env, mark the input optional, or drop it from the spec |
AdapterResolutionError: env contract carries no adapter tags |
resolving from a contract on an untagged env | serve with tags= or tag(); see Adapters |
AdapterResolutionError: ... has spec=None |
a tagged env paired with a model that declares no spec | pass spec=<ModelSpec>, or rlmesh.NO_ADAPTER to opt out of adaptation |
ValueError: Endpoint ... serves N environments |
RemoteEnv connected to a multi-env endpoint |
connect with RemoteVectorEnv instead (see Connect a Remote Environment) |
ConnectionError |
the env/model endpoint is unreachable or the connection dropped | confirm the address and that the server is up; re-dial a fresh client |
TimeoutError |
a connect or request exceeded its deadline | check the endpoint is serving; retry the operation |
EnvironmentException |
the env reported NotReady, Busy, Internal, Crashed, or Closed | reset the env, or inspect the env-side logs for the underlying fault |
RuntimeError: model error: ... |
a served model handler declined the request | check the model’s prediction code against the obs it actually receives |
ValueError: device=... requires a torch/jax ... |
device= on a numpy env/model |
drop device=, or set framework="torch" / "jax" |
ImportError: Failed to import _rlmesh native module |
the compiled extension is missing | reinstall the wheel for your platform |
Connection loss and crashes mid-episode
The eval loop does not retry a failed step. When the env connection drops or a served model fails while an episode is in flight, the call (predict or step) raises, and run() propagates it. The loop ends the open episode and releases the session before re-raising; it does not return a partial RunResult. A model instance or environment you passed remains yours to close.
run starts an episodeEpisodeResultRunResult- more episodes: back to run starts an episode
A few consequences follow from this.
A transport fault during an established run is reported as ConnectionError, not as a recoverable retry. RLMesh classifies transport conditions internally (a server that is still binding is treated as transient), but the public Python clients do not reconnect a dropped session for you. To recover, construct a fresh client and start a new run.
A served model handler that raises becomes a RuntimeError carrying the handler’s message. The env stays up, so you can fix the model and re-dial without restarting the env server.
Recovering cleanly
Wrap a run in a normal try/except and decide per family. Resolution errors are wiring bugs you fix in the spec or tags; transport errors call for a re-dial; an EnvironmentException usually means resetting or restarting the env.
import rlmesh
import rlmesh.adapters as adapt
try:
result = model.run(env, seeds=range(100))
except adapt.AdapterResolutionError as exc:
raise SystemExit(f"adapter mismatch, fix the spec or tags: {exc}")
except (ConnectionError, TimeoutError) as exc:
... # re-dial a fresh client and retry the run
except rlmesh.EnvironmentException as exc:
... # the env reported a fault; reset or restart it
Use session() as a context manager to release resources created for that session, including a sandbox model container. A model or environment instance you supplied remains yours to close.
Debugging a misaligned adapter
If a model performs unexpectedly on another environment, check the adapter before interpreting the score. A wrong image layout, rotation convention, or state scale can produce valid values with the wrong meaning.
The three tools run from static to live: adapter.explain() prints the exact transforms the resolver chose; the read / reader API extracts any role from a raw observation in whatever encoding you ask for; and join advisories warn you, at authoring time, about a tag that looks mis-declared. Start with explain(), reach for read when you need to see the values, and lean on advisories to catch mistakes before a peer ever connects.
Read what the resolver chose
resolve() returns an Adapter, and adapter.explain() prints every transform it derived: each resize, layout transpose, encoding conversion, range map, key remap, slice, and clip. Call it once, before you run a step.
import rlmesh.adapters as adapt
adapter = adapt.resolve(tags, env.observation_space, env.action_space, spec)
print(adapter.explain())
For the manipulation pair in Adapters, the output shows the primary image resized, the rotation going quat_xyzw -> rot6d, the instruction key remapped (goal -> task), and the model’s 6-D action rotation converted back to the env’s 3-D axis_angle and clipped. If a transform you expected is missing, or one you did not intend is present, the spec and the tags disagree about that leaf. When the spec uses a CustomEncoding, explain() also lists the host-side repack arms under a host-side encodings: section.
On the served path, the same description is available from the resolved adapter the model server builds, and from the contract a client receives. The point is the same either way: read the description, not the success rate.
Inspect observations by role
explain() tells you the plan. The read / reader API tells you the values. Both give a read-only, role-addressed view of a raw observation through the same adapter pipeline a model uses, so they are encoding-agnostic across envs and never mutate the observation.
sess.reader(*items) resolves once and returns a callable mapping a raw observation to a {role: value} dict:
import rlmesh.adapters as adapt
with model.session(env) as sess:
read = sess.reader(adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc"), adapt.EEF_POS)
obs, _ = sess.reset(seed=0)
view = read(obs)
print(view[adapt.IMAGE_PRIMARY].shape, view[adapt.IMAGE_PRIMARY].dtype)
print(view[adapt.EEF_POS])
sess.read(obs, item) is the one-shot single-role form; the underlying reader is cached per item, so calling it every step does not re-resolve:
ee = sess.read(obs, adapt.EEF_POS)
img = sess.read(obs, adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc"))
An item is one of two things, and the difference is exactly what you want when debugging:
- A bare role constant (
adapt.EEF_POS,adapt.IMAGE_PRIMARY) returns the quantity in the env’s own declared encoding and layout. Use it to confirm what the camera actually sends, before any conversion. - A model-input leaf that declares an encoding (
adapt.Image(adapt.IMAGE_PRIMARY, layout="hwc"),adapt.State(adapt.EEF_ROT, encoding="rot6d")) returns the value converted to that form. Use it to confirm the conversion produces what the model expects.
Reading the same role both ways, bare and as a leaf, shows you the transform end to end. The env must publish adapter tags (via an EnvFactory or tag()); otherwise there are no roles to address and the read raises AdapterResolutionError. Values come back in the env’s own framework, NumPy for a Gymnasium env, torch for a torch route. The reader vocabulary and the read API are covered alongside the eval loop in Running Evaluations.
See it live
For a moving target, attach the built-in viewer with view= on session. It shows the env’s render() frame plus every declared camera role, selectable at runtime, with a step/reward HUD. (run() drives the native runtime loop, which has no viewer seam – the viewer is fed from the session’s step loop.)
with model.session(env, view="terminal") as sess: # half-block frames in the terminal
sess.run(seeds=range(5))
with model.session(env, view="http:9000") as sess: # serve frames over HTTP on :9000
sess.run(seeds=range(5))
The string shorthands are "terminal", "http", "http:PORT", and "both"; construct rlmesh.View(...) directly to tune fps, image format, or quality. The viewer is best-effort: any setup failure disables it with a warning and never breaks the eval. It displays the environment’s image roles. To inspect the model’s transformed input, use its declared image spec with read or inspect the adapter output.
Join advisories and warnings
Some mismatches are not resolution errors, they are smells: a layout that looks mis-declared, a camera the env never provides that a model wants zero-filled, an aspect crop that discards pixels. RLMesh surfaces these as non-fatal advisories at two points.
At authoring time, tag() runs the native join check against the env’s spaces and emits each advisory through Python’s warnings, so you see your own mistake when you tag the env rather than in a peer’s serve logs.
import warnings
import rlmesh.adapters as adapt
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
env = adapt.tag(env, tags)
for w in caught:
print(w.message) # e.g. a layout hint for a frame that looks CHW
At resolve time, an adapter that fabricates or drops data records it. adapter.advisories() returns that subset of the description, the per-env data-loss and fabrication notes (a zero-filled camera, an aspect crop), empty when nothing noteworthy happened. Each note carries a severity: "info" for benign hints and explicitly-requested lossy steps, "caution" when the adapter substituted or fabricated data the model consumes (a camera bound to a different role than the model asked for, a zero-filled frame). A managed runner can log everything and choose to hard-fail on the caution tier alone.
for note in adapter.advisories():
print(f"[{note.severity}] {note}")
if any(note.severity == "caution" for note in adapter.advisories()):
raise RuntimeError("the model is not seeing what its spec asked for")
Symptom, cause, fix: a run that resolved but behaves wrong
These mirror the resolver pitfalls in Adapter Reference, framed for debugging a run that resolved cleanly but behaves wrong.
| Symptom | Cause | Confirm it | Fix |
|---|---|---|---|
| Image looks scrambled or rotated 90° | HWC vs CHW layout mismatch | read the role bare, then as Image(role, layout="hwc"), compare |
set layout on the model Image to what it wants |
| Wrong channel count slips through | RGB vs grayscale never declared | check read(obs, Image(role)).shape[-1] |
set channels on the Image to make a mismatch error |
| Policy acts as if the arm is mis-oriented | quat_xyzw vs quat_wxyz, or wrong base |
read the rotation role bare to see the env’s packing |
match the env’s exact encoding on the model side |
| Proprio values out of the expected range | scale mismatch between env and model | read the role as a State leaf and inspect the magnitudes |
set range on the model side to map it |
| A camera frame is all black | env lacks the role, filled by optional |
adapter.advisories() lists the zero-filled camera |
provide the camera, or accept the fill deliberately |
| Image edges or aspect look cropped | fit="crop" chose to discard pixels |
adapter.advisories() lists the crop |
use fit="pad", or match the target aspect |
explain() omits a transform you expected |
the spec and tags disagree on that leaf’s role | read explain() line by line against your spec |
align the role/encoding on whichever side is wrong |
Read raises AdapterResolutionError |
the env publishes no adapter tags | check env.metadata for the env-tags key |
serve with tags= or tag() |