Bring-Your-Own Container

Build and test EnvFactory and model containers for local and managed evaluations.

Start with an EnvFactory and serve it through python -m rlmesh.serve. The same image can run locally and on RLMesh Managed. This example checks the complete evaluation path with Pendulum and a policy that always applies zero torque.

Create two directories, env/ and model/. The files below are a complete runnable example.

Environment container

Save this as env/environment.py:

import gymnasium as gym
from rlmesh import EnvFactory
from rlmesh.adapters import Action, Actuator, EnvTags, StateTag


class Pendulum(EnvFactory):
    tags = EnvTags(
        observation=StateTag("x/pendulum_state"),
        action=Action(Actuator("x/pendulum_torque", dim=1, range=(-2.0, 2.0))),
    )

    def make(self) -> gym.Env:
        return gym.make("Pendulum-v1")

make() constructs each environment. Optional prepare() handles one-time setup, and tags publishes the observation and action roles for model adapters. The x/ roles here are specific to this example; both sides agree on their meaning.

Save this as env/Dockerfile:

FROM python:3.11-slim
ARG RLMESH_VERSION=0.1.0
RUN pip install --no-cache-dir "rlmesh[gymnasium,numpy]==${RLMESH_VERSION}"
WORKDIR /app
COPY environment.py .
EXPOSE 50051
CMD ["python", "-m", "rlmesh.serve", "--env", "environment:Pendulum"]

Model container

Save this as model/model.py:

import numpy as np
from rlmesh.adapters import Action, Actuator, ModelSpec, State
from rlmesh.numpy import Model


class ZeroTorque(Model):
    spec = ModelSpec(
        input=State("x/pendulum_state", dim=3),
        output=Action(Actuator("x/pendulum_torque", dim=1, range=(-2.0, 2.0))),
    )

    def predict(self, observation):
        return np.zeros(1, dtype=np.float32)

The ModelSpec describes the three observation values and one torque output. It lets the managed probe construct test inputs without a running environment. Returning a float32 array with shape (1,) matches Pendulum’s continuous action space.

Save this as model/Dockerfile:

FROM python:3.11-slim
ARG RLMESH_VERSION=0.1.0
RUN pip install --no-cache-dir "rlmesh[numpy]==${RLMESH_VERSION}"
WORKDIR /app
COPY model.py .
EXPOSE 50051
CMD ["python", "-m", "rlmesh.serve", "model:ZeroTorque"]

Check and build

Install the SDK locally and check both classes before building. rlmesh check and rlmesh check-image are available from the SDK release that includes them; until then skip the check commands. rlmesh check fails on the packaging mistakes the platform would reject at push: a model without spec, an environment without tags, or a spec that does not resolve.

pip install "rlmesh[gymnasium,numpy]==0.1.0"
(cd env && rlmesh check environment:Pendulum)
(cd model && rlmesh check model:ZeroTorque)

From the parent directory, build both images for the managed fleet’s architecture, then check the built images:

docker build --platform linux/amd64 -t my-env:latest env
docker build --platform linux/amd64 -t my-model:latest model
rlmesh check-image my-env:latest
rlmesh check-image my-model:latest

check-image reads the serve command, the exposed port, the architecture, and the dev.rlmesh.* labels off the image config. Fix anything it reports as failed; warnings do not block a push.

Test locally

Run each image in its own terminal:

docker run --rm -p 127.0.0.1:50051:50051 my-env:latest
docker run --rm -p 127.0.0.1:50052:50051 my-model:latest

Both containers listen on port 50051 internally. rlmesh.serve respects RLMESH_ADDRESS, which the platform assigns when it runs the images. Published ports are bound to loopback because these endpoints are plaintext and unauthenticated; see Exposure and authentication.

Run this client in a third terminal:

import rlmesh
from rlmesh.numpy import RemoteEnv, RemoteModel

env = RemoteEnv("127.0.0.1:50051")
try:
    with rlmesh.session(RemoteModel("127.0.0.1:50052"), env) as sess:
        obs, _ = sess.reset(seed=0)
        total_reward = 0.0
        steps = 0
        while not sess.done:
            obs, reward, terminated, truncated, _ = sess.step(sess.predict(obs))
            total_reward += reward
            steps += 1
        print(f"Episode: {steps} steps, reward {total_reward:.2f}")
finally:
    env.close()

Expect 200 steps and a negative reward. This exercises discovery of the environment contract, adapter resolution, remote model prediction, and the reset/step loop.

Run on RLMesh Managed

Follow Upload a Model or Environment to log in and find your organization namespace. Tag and push the same images:

docker tag my-env:latest registry.rlmesh.dev/your-namespace/pendulum:latest
docker tag my-model:latest registry.rlmesh.dev/your-namespace/zero-torque:latest
docker push registry.rlmesh.dev/your-namespace/pendulum:latest
docker push registry.rlmesh.dev/your-namespace/zero-torque:latest

Use the registry host shown in your dashboard. Wait for both images to pass their probes, then select them in New evaluation. A local episode verifies the serving path; the managed probes additionally validate fleet compatibility and runtime behavior.

The standard rlmesh.serve command lets the platform identify each image and read its description from the handshake. A custom entrypoint needs a dev.rlmesh.describe label as well as the same serving behavior. Changing only the model’s Dockerfile command is insufficient if it has no ModelSpec.

Neither image here needs a GPU. For one that does, add a gpu tag on the repository in the dashboard, or bake it into the image: LABEL dev.rlmesh.package='{"schemaVersion":1,"name":"my-model","tags":["gpu"]}'. Without a tag the probe infers a GPU from a CUDA base image and confirms it by measuring VRAM use; a cpu tag forces CPU. See GPU images.

Version pinning

These Dockerfiles use rlmesh==0.1.0; set --build-arg RLMESH_VERSION=<version> to build on another release, and use the version your managed deployment pins. From the SDK release that advertises workflow editions in rlmesh describe, the platform admits an image when it shares a workflow edition with the runtime instead of requiring the same SDK version: every release carries the sealed edition it implements (for example 2026.06) beside its own build, so a platform release does not invalidate a pushed image unless the sealed edition changes, and the push probe reports the negotiated edition or why none was shared. See Compatibility.

Optional sandbox workflow

For local container lifecycle management, replace RemoteModel(...) in the client with rlmesh.SandboxModel("image://my-model:latest"). Closing the session then stops the model container it started. SandboxModel is experimental and does not upload an image or launch a managed evaluation. See Sandbox Examples for source-built environment sandboxes.