Custom Environment and Model

Serve a small Python environment and run a model against it.

An EnvFactory can return a plain Python object with Gymnasium-style methods. This example serves a tiny counter, runs a model against it, and then drives the loop by hand. It uses the NumPy backend throughout. For the full authoring surface, see Environments and Models.

Serve a custom environment

Save this as counter.py. CounterEnv implements the environment behavior; the Counter factory constructs it for local evaluation or serving.

import rlmesh


class CounterEnv:
    observation_space = rlmesh.spaces.Discrete(5)
    action_space = rlmesh.spaces.Discrete(2)

    def __init__(self):
        self.step_count = 0

    def reset(self, seed=None, options=None):
        self.step_count = 0
        return 0, {}

    def step(self, action):
        self.step_count += 1
        observation = self.step_count % 5
        terminated = self.step_count >= 3
        return observation, 1.0, terminated, False, {"action": action}

    def close(self):
        pass


class Counter(rlmesh.EnvFactory):
    def make(self):
        return CounterEnv()

Any object with the same shape serves the same way: an observation_space, an action_space, reset(seed=None, options=None), step(action), and close(). Run it:

python -m rlmesh.serve --env counter:Counter --address 127.0.0.1:5555

Run a custom model

Save this as evaluate_counter.py. The model’s predict method takes an observation and returns an action; run evaluates complete episodes against the served endpoint.

from rlmesh.numpy import Model


class CounterPolicy(Model):
    def predict(self, observation):
        return 0


if __name__ == "__main__":
    result = CounterPolicy().run("127.0.0.1:5555", episodes=1)
    print(f"episodes={result.num_episodes} mean_reward={result.mean_reward:.2f}")

Run it against the server:

python evaluate_counter.py

Drive the loop yourself

To step the environment by hand instead of handing it to a Model, examples/python/quickstart/eval.py opens a RemoteEnv and runs a sampled-action loop.

from rlmesh.numpy import RemoteEnv

env = RemoteEnv("127.0.0.1:5555")
obs, info = env.reset(seed=0)
for step in range(1, 65):
    action = env.action_space.sample()
    obs, reward, term, trunc, info = env.step(action)
    if term or trunc:
        break
env.close()

The server owns the environment. The model, or your own loop, connects to its address and returns actions.