Quickstart
Serve CartPole, connect a client, and run a model locally.
In two terminals, you’ll serve CartPole, connect a client, then run a model against the same endpoint. You will define an EnvFactory and a Model subclass, the same authoring pattern used by catalog images. The example policy always chooses action 0 to keep the evaluation loop easy to inspect.
Install
pip install "rlmesh[gymnasium,numpy]"
See Installation for the optional extras.
Server
Save this as server.py, then run python server.py in the first terminal and leave it running:
import gymnasium as gym
import rlmesh
class CartPole(rlmesh.EnvFactory):
def make(self):
return gym.make("CartPole-v1")
if __name__ == "__main__":
CartPole().serve("127.0.0.1:5555")
The EnvFactory class is also the entrypoint for a container: python -m rlmesh.serve --env server:CartPole. Keep serving behind the __main__ guard so importing the class does not start a second server. See Author an Environment for tags, setup, and construction parameters.
Client
Save this as client.py, then run python client.py in a second terminal:
from rlmesh.numpy import RemoteEnv
env = RemoteEnv("127.0.0.1:5555")
obs, info = env.reset(seed=0)
terminated = truncated = False
total_reward = 0.0
while not (terminated or truncated):
action = env.action_space.sample()
obs, reward, terminated, truncated, info = env.step(action)
total_reward += reward
print(f"Episode reward: {total_reward}")
env.close()
The second terminal prints an episode reward when CartPole finishes. The server owns the Gymnasium environment and its dependencies. The client learns the action and observation spaces when it connects, so it does not import CartPole.
Run a Model Against It
Save the following as evaluate.py and run python evaluate.py in the second terminal. A Model subclass keeps the policy in an importable class that you can also serve or package:
from rlmesh.numpy import Model
class Policy(Model):
def predict(self, observation):
return 0
if __name__ == "__main__":
result = Policy().run("127.0.0.1:5555", episodes=3)
print(result.mean_reward)
run connects to the same endpoint, plays three episodes, and returns a RunResult with per-episode rewards and mean_reward. success_rate is None here because CartPole does not report a task-success flag. The reward will vary because this example uses a fixed action rather than a trained policy.
The same classes work in one process: import CartPole from server and Policy from evaluate, then call Policy().run(CartPole(), episodes=3). Add weight loading in Policy.load() when you replace the fixed-action policy. For a managed image, add the environment tags and model spec shown in Bring Your Own Container, then use rlmesh.serve as the container command.
To run the repository’s example files or swap environments by flag, follow the CartPole walkthrough.
Next
- Running Evaluations:
run,session, seeds, and reading results. - Models: subclassing a backend
Model, loading weights, and defining its spec. - Upload a Model or Environment: bring a locally tested container into the managed catalog.