Performance and Scaling
Batching, action chunks, payload size, and device placement.
Evaluation throughput depends on how many model calls you make, how much data each call carries, and how many environment lanes can run at once. This page covers the available batching, action-chunk, and payload choices. For startup, session lifetime, and connection failures, use the serving, evaluation, and troubleshooting guides linked below.
Batch the forward pass
A vectorized route runs many env lanes against one model. By default the runtime calls the model once per lane. Define predict_batch (alongside predict) and the runtime instead fuses the N per-lane observations into one batched observation and calls the model once for the whole vector.
from rlmesh.torch import Model
class Policy(Model):
def predict(self, obs):
return self._forward(obs[None])[0] # one lane
def predict_batch(self, obs):
return self._forward(obs) # N lanes in one forward
The fused observation gives every leaf a leading batch axis, so a Dict observation arrives as {key: array[N, ...]}, the shape an RL or VLA stack already hands a policy. Return the batched action the same way; the runtime splits it back per lane. The engine prefers this corner for a vectorized route and falls back to per-lane predict when it is absent.
For a policy that emits an action chunk, predict_chunk_batch is the batched chunk corner; the runtime splits the batch axis per lane and replays each chunk. The four corners and how the runtime derives the ones you skip are in Models.
Fewer model calls per episode
A model that predicts an action chunk can execute several actions before re-planning. Pass execution_horizon to run or session, and the runtime calls the model once, executes that many actions of the returned chunk one per env step, then calls again.
result = model.run(env, seeds=range(50), execution_horizon=8)
The runtime executes up to execution_horizon actions per model call, bounded by the returned chunk length and the episode boundary. This reduces prediction calls while the policy acts open-loop between predictions. A model without a chunk prediction method warns and runs one prediction per step. See Running Evaluations.
If the model produces a fixed chunk length K, declare native_chunk = K and return the whole chunk. RLMesh rejects an execution horizon above K during adapter resolution and rejects predictions whose length differs from the declaration.
Overlap inference with replay
The loop above is synchronous: the runtime calls the model when the chunk runs out, and the env waits for the forward. Pass prefetch_lead to run and the runtime predicts the next chunk while prefetch_lead (or fewer) replay frames of the current one still execute, so the forward overlaps the env steps instead of stalling them.
result = model.run(env, seeds=range(50), execution_horizon=8, prefetch_lead=2)
The prefetched chunk is conditioned on an observation up to prefetch_lead steps stale. That is how a deployed asynchronous policy behaves, and it is a different measurement: scores from a run with a lead are not comparable to the synchronous loop, so label them. A chunk prefetched across an episode boundary is discarded, and the new episode re-plans from its own reset observation. A lead at or above the chunk length asks for the next chunk as soon as the current one starts playing. The default of 0 is the synchronous loop. The lead is a native-loop knob: session steps in Python and has no replay to overlap, and a served model driven through rlmesh.run rejects it.
Frame-stack overhead
A model that conditions on a short history declares stack=N on an image input. The env still sends one frame per step; RLMesh buffers the last N processed frames and emits them on a new leading axis.
"image": adapt.Image(adapt.IMAGE_PRIMARY, size=224, stack=4)
The cost is local at execution_horizon=1. On the in-process run path the buffer is a host-side rolling deque of N processed frames per stacked input, cleared on reset; on the served path the core keeps the same buffer, keyed per episode. Nothing extra crosses the wire, so stacking trades a little host memory and a stack copy for keeping the network payload at one frame per step. The buffer holds the processed (resized, converted) frame, so its size follows the model’s target resolution, not the camera’s.
Above execution_horizon=1 the window still has to see every step, including the ones the runtime executed from a replayed chunk without predicting. A Python session ticks the window itself. The native runtime and a served model negotiate it when the adapter resolves (the model answers which inputs hold a window), and the runtime then carries each replayed step’s raw observation to the model as history rows on the next predict, stamped with a step counter the model holds to consecutive values. So a stacked model at horizon H pays H - 1 extra observations of wire bytes per re-plan, and the history.rows metric on the run’s telemetry says how many rode each predict. Async prefetch still works on such a route: the prefetched request carries the rows buffered so far, and the observation it predicts from is its own rather than a row, so every step still reaches the model once. A route whose windows would hold more than RLMESH_FRAME_HISTORY_LIMIT_BYTES (default 2 GiB) is refused when the adapter resolves, and the same budget bounds the whole model endpoint: a predict that would open windows for a new lane is refused while every route’s live lanes, at their full window size, would exceed it.
A model that reads its own last command (adapt.Previous(adapt.ACTION_JOINT_POS)) is in the same class. Its window is one raw action row per lane rather than frames, so the memory is negligible, but every executed step has to be numbered so the row for the step before a re-plan is the frame the runtime actually replayed. The route therefore negotiates observation history like a stacked one: replayed steps ride as history rows on the next predict (whole observations, so the same H - 1 observations of wire bytes per re-plan), the model holds them to consecutive steps, and async prefetch carries the rows the same way.
Payload encoding
The adapter encodes only the observation keys the plan actually reads. An env that returns extra keys, or one unencodable key, does not pay for them and does not abort a step over them. This is automatic; there is no flag, and explain() shows which keys the plan touches (see Troubleshooting).
Values travel as framework-neutral bytes directed by the spec, so the env’s framework and the model’s framework are independent and neither forces a conversion on the other. The smaller you make the model’s declared input (a single primary camera instead of three, a target resolution the policy needs rather than the camera’s native one), the less there is to encode and move. That is a spec decision, made once.
One encoded message is capped at 256 MiB. RLMesh raises tonic’s 4 MiB default to that on every env and model client and server, in both directions, and the Python SDK inherits it; the cap is a build constant, not a ServeOptions knob. It bounds a single protobuf message (one step response, one predict request, one grouped predict batch), not a stream or a run, so it binds on the widest single step: num_envs times one observation’s encoded bytes. An oversized message never reaches the peer’s handler – the sender fails its own encode and a receiver aborts its decode, both reporting the gRPC status OUT_OF_RANGE with a message length too large detail naming the found length and the limit, which surfaces as a transport error rather than an env or model error. If you hit it, step fewer lanes at once or shrink the declared observation (a smaller dtype, one camera, a lower target resolution); there is no setting that raises it.
Device and framework placement
For a torch or JAX model, set device in load alongside moving your weights. RLMesh moves every observation tensor leaf onto that device before predict, so you never call .to(device) in the hot path and there is one source of truth for placement.
class Policy(Model):
device = "cuda:0"
def load(self):
self.weights = load_checkpoint().to(self.device)
On the env side, EnvServer takes framework= ("torch" / "jax" / "numpy") to type the obs/action seam, and device= to place the incoming action for a torch/jax env. device= requires a framework with a device; a numpy env or the default backend rejects it. Observations need no declaration, a torch/jax obs (GPU included) is auto-detected and encoded either way. See Serve an Environment.
Vector endpoints
One endpoint can serve many env instances. EnvServer detects a vectorized env (one exposing num_envs and single_* spaces) and serves a vector endpoint automatically; connect to it with RemoteVectorEnv.
import gymnasium as gym
import rlmesh
envs = gym.vector.SyncVectorEnv([lambda: gym.make("CartPole-v1") for _ in range(4)])
rlmesh.EnvServer(envs, "127.0.0.1:5555").serve()
A vector endpoint plus a batched model corner is the fast combination: N lanes step together and one model forward covers them. The client side is in Connect a Remote Environment.
A list of scalar envs (EnvServer([make() for _ in range(N)]), or RLMESH_NUM_ENVS=N on a prebuilt image) serves N lanes instead: the endpoint advertises the subset_step capability, and a runtime drives each lane as its own episode loop over the one Join stream, keeping a step or reset per lane in flight. Lanes finish and restart episodes independently, and each episode’s seed and index are fixed by a route-global slot counter, so the scored set is identical however the lanes’ timing interleaves. Batched model corners still apply: concurrent lane predicts are grouped at the model. Prefer lanes over a gym vector env whenever episodes have uneven lengths or resets are slow (scene rebuilds, GPU renderers).
What is not a knob
Some things that look tunable are fixed, automatic, or not exposed in the Python API. Knowing which is which saves you looking for a setting that is not there.
| Looks like a knob | Reality |
|---|---|
Per-step request timeout inside run / session |
Not a loop-level knob. Pass request_timeout_seconds= to RemoteEnv / RemoteVectorEnv to bound each call. |
| Client-side retry / reconnect | Not done for you. A dropped session raises; re-dial a fresh client (see Troubleshooting). |
predict_concurrency (server pipelining) |
Present in the Rust serve options but not in the Python ServeOptions constructor. |
| Adapter conversions (resize, normalize, encoding) | Correctness transforms the resolver chooses, not performance dials. Shrink the spec to do less work. |
| Frame-stack depth | A model capability set by stack=N, sized by the model’s needs, not a tuning parameter. |
The lifecycle options that ServeOptions does expose, idle_timeout_seconds, drain_timeout_seconds, close_timeout_seconds, and allow_remote_shutdown, control shutdown behavior rather than throughput; they belong to the session lifecycle covered below.
Long-running evaluations
Wait for the environment endpoint to be ready before connecting; the Python serve CLI exposes --ready-fd for this. Serve an Environment covers startup and readiness signals.
Use model.run(..., hooks=...) or Session.run(..., hooks=...) to observe episode progress, and RunResult to read final counts and timing. Use a session() context manager when you need manual step control so the connection closes on exit. The Evaluation guide covers both paths and episode accounting.
A disconnected session raises an error; create a new client for the next attempt. See Troubleshooting for transport failures and recovery.