Edition 2026.06
This document is the behavioral contract identified by the workflow edition base 2026.06.
- Status: sealed as
2026.06at v0.1.0 - Protocol generation:
rlmesh-wire-v1
This document defines the behavior of a session running the 2026.06 workflow edition. The runtime, environment, and model follow these semantics for the rest of that session.
The protobuf files for rlmesh-wire-v1 define the wire shape; this document defines what conforming use of that shape means. Where an implementation and this document disagree, the implementation has a bug. A change that alters the meaning of an interaction described here mints a new edition; it does not amend this one once sealed.
Session Establishment
A session is one successful Handshake followed by one Join stream. Conforming clients do not open Join before completing a handshake with compatible = true.
- The client states its wire protocol in
protocol_generation, lists every edition it can operate under insupported_workflow_editions, and declares its preferred edition inpreferred_workflow_edition. The server replies with its own supported and preferred editions. The handshake does not select a session edition. - The server replies
compatible = truewhen the protocol generations are compatible. Generation is the only handshake-level decision. Only the runtime sees every participant, so it selects the session edition across env, model, and runtime and records it onResolveAdapterandConfigureEnv. A generation-compatible peer that shares no edition handshakes successfully, then fails at the runtime’s edition selection. - On
compatible = false,error_messageexplains the (generation) failure andsupported_workflow_editionslists the server’s editions for diagnostics. The session ends; there is no renegotiation round. - Capability maps are advisory in both directions: a key mapped to the literal value
"true"means the named feature is available; an absent key, or a key with any other value, means it is not (other values are reserved for a future value grammar). Capabilities gate optional features: a capability may scope which lanes an interaction covers (see Partial width) or how a server schedules its responses, but it never changes what an interaction means for the lanes it covers. - The environment handshake response carries the
EnvContract(spaces, metadata, render mode,num_envs) whenevercompatible = true. The contract is fixed for the life of the session.
Environment Workflow (rlmesh.env.v1.EnvService)
Ordering
Requests on a Join stream are processed strictly in arrival order, one at a time, except for the lane-scoped carve-out below. Every request produces exactly one response, carrying the originating request_id. There is no reordering and no silent dropping: a request the server cannot satisfy produces an in-band EnvError response.
A server that advertises the subset_step capability at handshake may process lane-scoped requests — a Reset or Step naming env_indices — concurrently and emit their responses in completion order. This is a scheduling choice (out of scope; see below), never a wire change: every response still mirrors its request_id, so a client demuxes by id. Ordering is not preserved between lane-scoped requests, so a client keeps at most one lane-scoped request in flight per lane; whole-vector ops (full-width Reset/Step, Render, Configure, Close) are processed only after in-flight lane ops drain, so they always observe quiescent lanes.
Vectorization
The served environment is a fixed-width vector of num_envs sub-environments, established by the handshake contract. Observations and actions are batched SpaceValues covering the sub-environments the request names — all of them unless it names env_indices; rewards, terminated_mask, and truncated_mask carry one entry per covered sub-environment, in index order (mask bytes are per-sub-environment flags; nonzero means set).
Partial width (subset_step). An env advertising the subset_step capability at handshake accepts a non-empty ResetRequest.env_indices / StepRequest.env_indices naming a subset of lanes. When non-empty, the batched action and observation rows, seeds, episode_ids, rewards, terminated_mask, and truncated_mask align positionally to those k indices rather than to all num_envs sub-environments, and StepResponse.env_indices echoes the lanes the response covers (empty means full width). completed_episodes is never positional: each EpisodeMetadata carries its own env_index. An env that does not advertise the capability rejects a non-empty env_indices with UNSUPPORTED.
Reset
Reset (re)starts all sub-environments and must precede the first Step of a session; a Reset naming env_indices restarts only those lanes and leaves the other lanes’ episodes running (see Partial width). seeds is either empty (server defaults apply) or carries one seed per lane the reset covers. ResetRequest.episode_ids carries the authoritative episode ids the runtime minted for the lanes being started; the env adopts them (it never mints) and tags its tracked episodes with them. The ResetResponse carries the initial batched observation and infos; it has no episode-id fields, since ids are reported back on Step in completed_episodes.
ResetRequest.options is an open map, except for a closed set of reserved keys the edition defines. This edition reserves exactly one, trial_index: the 0-based ordinal of the episode being started, an integer for a single lane or a per-lane list aligned to the lanes being reset, walked in order so a run covers a benchmark’s trials once each. A reserved key is delivered only to an env whose contract metadata declares it under rlmesh.env.v1.reset_options (a list of key names); an env that declares nothing receives no reserved key, so forwarding options blindly into a third-party reset stays safe. Names outside the reserved set are caller-defined and untouched by the runtime.
Step
Step applies one batched action to the sub-environments the request covers (all of them, or those named in env_indices). StepRequest.episode_ids carries the runtime’s authoritative per-lane ids (the env adopts a rolled id when it autoresets a lane under NEXT_STEP). The response carries the next batched observation, per-sub-environment rewards and termination/truncation masks, shared infos, and completed_episodes metadata for episodes that ended on this step. The env reports lifecycle via completed_episodes, and each EpisodeMetadata carries the runtime-minted episode_id (alongside its env_index) that the env adopted at the episode’s start; the runtime remains the sole id authority, the env only echoes the ids it was handed.
Episode Accounting
A tracked episode per sub-environment begins at Reset. Each Step accrues to it. When a sub-environment reports terminated or truncated, its episode completes: metadata (id, seed, step count, cumulative reward, termination cause, timing, final_info) is delivered once in that response’s completed_episodes. The id in that metadata is the runtime-minted id the env adopted at the episode’s start.
The edition itself never restarts a sub-environment: it sends no Reset of its own. A lane’s restart behavior is fixed for the session by EnvContract.autoreset_mode, and under NEXT_STEP an autoreset roll establishes a new tracked episode without a Reset.
Autoreset (EnvContract.autoreset_mode).
AUTORESET_MODE_NEXT_STEP: the stepton which a lane reports terminated or truncated carries that episode’s final observation, and itsEpisodeMetadataappears in that response’scompleted_episodes. The env restarts that lane on the nextStep(t+1): it adopts the id the runtime supplies for that lane in that step’sepisode_ids, does not apply the action supplied for that lane, and returns the new episode’s initial observation with reward 0 and both masks clear for that lane. That roll begins a new tracked episode at step 0. A serving side that reports a terminal or a nonzero reward for a lane on its roll step, or that steps a lane with no active episode and no pending roll, is a non-recoverableInternalerror.AUTORESET_MODE_DISABLED: only an explicitResetrestarts a lane.AUTORESET_MODE_UNSPECIFIED: treated asDISABLED.AUTORESET_MODE_SAME_STEP: reserved on the wire and not implemented in this edition; a session declaring it is refused.
Render and Close
Render returns a PNG frame when the environment supports the contract’s render mode, and no frame otherwise. Close ends the session: the server responds once (with metadata for episodes it finalizes) and then closes the Join stream. No further requests on that stream are answered.
Timeouts and Errors
A positive timeout_ms on a request is a server-enforced deadline; expiry produces an in-band EnvError with code TIMEOUT. Errors are reported in-band as EnvError responses with a code, message, and is_recoverable: a recoverable error leaves the session usable for further requests; a non-recoverable error means the client must abandon the session.
Shutdown
Shutdown is a unary request to terminate the endpoint itself, distinct from closing a session. The server may refuse it (accepted = false), and endpoints may disable remote shutdown entirely.
Model Workflow (rlmesh.model.v1.ModelService)
The model service reverses the dialing direction: the runtime connects to a served model participant. Session establishment is as above (without an environment contract in the response).
Adapter Lifecycle
The model-facing protocol is keyed by two ids: env_id (a connected env container, UUIDv7, minted by the runtime) and episode_id (one rollout, UUIDv7, minted fresh on every reset). There is no positional lane/slot on the wire. An adapter is the resolved model-side context for one env_id (one adapter per env, no sharing), carried in AdapterContext.
Five ops, all keyed by env_id:
ResolveAdaptermust precede the firstPredictfor an env and fixes that env’sEnvContract(idempotent upsert). APredicton an unresolved env produces aModelErrorwith codeNOT_CONFIGURED.Predictinfers through the resolved adapter (below).GroupedPredictcarries N independentPredictrequests in one message, each already routed to its own resolvedenv_id; the response carries one result per group in request order, and a group that fails reports its ownModelErrorwithout sinking the others.ResetAdapterdrops per-episode adapter state (frame-stack buffers): withepisode_idsset, exactly those keys; empty, all of the env’s episode state. The adapter stays resolved. Because UUIDv7 ids never repeat, a missed/lateResetAdapteronly leaks memory; it can never alias a new episode, so it is GC, not a correctness gate. It is also the episode-END edge the model sees: withepisode_idsset, the model’s own end hook fires once per listed id, naming it; empty is an evict-all/teardown and fires nothing.ReleaseAdapterremoves the adapter entirely (implies reset-all).Closerequests graceful shutdown of the whole participant.
Predict
A Predict request carries the AdapterContext (env_id, session_id, request_id), a batched observation encoded per the env’s observation space, and an ordered episode_info vector, the self-describing batch: row i of the observation belongs to episode_info[i], decided per tick. The model keys all per-episode state by episode_id, never by position-across-ticks; it lazy-seeds an episode’s state on its first appearance and evicts it on ResetAdapter. episode_info is per row for every shape of forward: a model that fuses N rows into one batched call receives the N identities row-aligned with the observation, never one identity for the call, because those rows may belong to independent episodes. The response mirrors the request’s context and carries an actions list whose first frame (actions[0]) is this step’s action batch (one row per episode_id), inheriting the request’s episode_info order.
actions[1..], when present, are pre-split future-step action frames for open-loop action chunking. The runtime buffers them and applies one per step without re-calling the model, then re-plans when the buffer drains or any lane’s episode rolls. This is scheduling data (out of scope; see below), not a change to the per-step contract: actions[0] still carries exactly one action row per episode_id, and a chunk-unaware runtime that uses only actions[0] re-plans every step and stays correct. How many frames a model emits is governed by execution_horizon, which the runtime pins on ResolveAdapter (the replay horizon h; 0/1 = no chunking). The horizon is a runtime decision, not part of the model spec.
Native Chunk
ResolveAdapterResponse carries optional uint32 native_chunk = 1 — the model’s answer for how many per-step actions ONE chunk corner call returns — and ObservationHistoryNeeds history = 2, the frame-window inputs a history-capable model answers (below). Absence of native_chunk means undeclared, never zero: the runtime then executes min(len(chunk), execution_horizon) of whatever comes back.
execution_horizon (h) is the runtime’s scheduling decision, pinned on ResolveAdapter. The native chunk (K) is a property of the model’s weights, so the model is the only participant that can state it, and it states it here rather than on the wire per predict. A model that declares K is held to it: a resolve at h > K is refused, and a Predict whose chunk is not exactly K frames is a ModelError. actions[0] still carries exactly one action row per episode_id either way, so a chunk-unaware runtime is unaffected.
Ordering and Errors
Every request on a Join stream is answered exactly once, mirroring request_id; failures are in-band ModelError responses, and is_recoverable has the same meaning as on the environment service.
A served model may accept observation history: the runtime declares on ResolveAdapter that it will deliver every env step it executes from a replayed chunk (delivers_history), the model answers the inputs that hold a frame window (ResolveAdapterResponse.history), and each later Predict carries those steps as history rows, oldest first, with a per-group step counter on rows and requests that the model holds to consecutive values. A model that does so advertises the rlmesh.model.observation_history.v1 capability at handshake.
A served model may pipeline requests: process them concurrently and emit responses in completion order rather than strict arrival order. This is a scheduling choice (out of scope for the edition; see below), never a wire change. Responses still mirror request_id, so a client demuxes them by id. A server that pipelines advertises the rlmesh.model.concurrent_predict.v1 capability at handshake; its absence means responses arrive in arrival order. Per-env ordering is always preserved: for a given env_id, the model applies ResolveAdapter, Predict, ResetAdapter, and ReleaseAdapter in the order the client sent them. A whole-session Close drains after every outstanding request. Requests for different envs may complete in either order.
Value Conformance
When the session edition is 2026.06, both peers agree on how observation and action values are checked against the spaces declared in the EnvContract. A value’s dtype is always coerced to its declared dtype before transport, so a delivered value’s dtype always equals the space the peer negotiated; a peer never receives a per-message dtype. A conformance warning may accompany a delivered value; a warning never withholds or alters the value beyond the coercion already applied.
Structural conformance
Regardless of the validation policy, a value is rejected when:
- a
Dictvalue is missing a declared key; - a
BoxorMultiDiscretevalue has the wrong rank or shape, aMultiBinaryvalue the wrong shape, or aTuplevalue the wrong arity; - a
Discreteelement is outside its domain, or aMultiDiscreteelement is outside its own per-element domain; - any element of a numeric value is
NaN.NaNis never a member of any space, so it is rejected even when other elements are merely out of bounds.
The recoverability of a structural rejection depends on which side produced the value. A rejected action is a recoverable EnvError (InvalidAction): the action is not delivered, but the session stays usable. A rejected observation is non-recoverable (Internal), because the serving side produced a value that violates its own contract. In both cases the offending value is never delivered.
Range conformance
For Box bounds and Text charset and length, the serving side carries a validation policy:
- warn (the default): the deviation is delivered and reported as a conformance warning.
- strict: the deviation is rejected, like a structural deviation.
- off: the range, charset, and length checks are skipped; structural conformance still applies.
A Box element is in range when it satisfies its declared bounds; an infinite or absent bound imposes no constraint on that side, so +/-inf is in range against a matching infinite bound. A Text value conforms when its length is within [min_length, max_length] (counted in characters) and, when the charset is non-empty, every character is in it. Observations and actions share one policy and both default to warn.
Enforcement
The serving side validates each observation it produces before transport and each action it receives before delivering it to the environment. Structural deviations are rejected regardless of policy; range deviations follow the policy.
Dtype conformance
A value is coerced to its declared dtype before transport:
- a value carried as a native RLMesh tensor must already have the declared dtype;
- a value supplied as a host array or sequence (for example NumPy) is coerced to the declared dtype, except that a float supplied for an integer dtype is rejected unless every element is finite, integral, and representable in the target integer dtype’s range, with no silent truncation and no out-of-range wraparound;
- a floating-point value supplied for a floating-point dtype is coerced to the declared dtype (a narrowing such as
float64tofloat32may lose precision).
Conformance warnings
A conformance warning is reported in-band under the reserved rlmesh.conformance.warning key in the info map returned by reset and step, at most once per (deviation kind, value path) per session. Conformance warnings are advisory and never make a session non-recoverable; the path format is advisory.
Errata
Clarifications and additive wire growth that do not change the meaning of any interaction above, and so do not mint a new edition. No errata as of the seal.
Out of Scope
This edition does not constrain transport security or authentication, client retry or reconnection policy, scheduling and batching strategy, performance, or the meaning of individual capability names. Those evolve freely without minting an edition.