rlmesh.adapters.Image

classfrom rlmesh.adapters import Image[source]

An image input expected by a model.

Image(
    role: str,
    height: int | None = None,
    width: int | None = None,
    layout: ImageLayout = 'hwc',
    channels: int | None = None,
    dtype: str = 'uint8',
    normalize: bool | tuple[float, float] = False,
    lead_dims: int = 0,
    upside_down: bool = False,
    resample: Resample = 'bilinear',
    allow_upscale: bool = False,
    fit: FitMode | Sequence[FitMode] | None = None,
    optional: bool = False,
    fill: int | None = None,
    stack: int = 1,
    size: InitVar[int | None] = None,
    *,
    crop: float | None = None,
    crop_area: float | None = None,
    crop_mode: CropMode = 'zoom',
    jpeg_quality: int | None = None,
    channel_order: ChannelOrder = 'rgb',
    render: int | tuple[int, int] | None = None,
    offsets: StackSpec | None = None,
    stack_pad: StackPad = 'first',
    part: str | None = None,
    stride: InitVar[int | None] = None,
)

There is no key – placement in the input tree is the payload position.

Attributes

role

attribute[source]
role: str

Semantic role matched against env image features.

height

attribute[source]
height: int | None

Target image height, or None to keep the env height.

width

attribute[source]
width: int | None

Target image width, or None to keep the env width.

layout

attribute[source]
layout: ImageLayout

Axis layout the model expects ("hwc" the default, or "chw"). This is the model author’s declaration: when it differs from the env’s layout the adapter transposes silently to match, so a model that wants "chw" must say so – an omitted layout means "hwc" and the env’s frame is fed through unchanged. (channels guards the channel count, not the axis order.)

channels

attribute[source]
channels: int | None

Channel count the model expects (e.g. 3 for RGB, 1 for grayscale). When set, a resolve error if the env image differs; the adapter does not convert between channel counts.

dtype

attribute[source]
dtype: str

NumPy dtype name the model expects.

normalize

attribute[source]
normalize: bool | tuple[float, float]

Whether (and into what range) to map 8-bit pixel values before casting: False (off, the default), True (the conventional [0, 1]), or a (low, high) pair for a model trained on a different range (e.g. (-1.0, 1.0)). One field, so an on/off flag can never disagree with a range; False is an authoritative off-switch.

lead_dims

attribute[source]
lead_dims: int

Number of leading singleton axes to add (batch/time).

upside_down

attribute[source]
upside_down: bool

Whether the model was trained on images rotated 180 degrees from the canonical upright orientation (a true rotation – rows and columns reversed – not a vertical flip). Declared on both ends: the adapter rotates only when the env and model disagree, so an env that already renders upside-down pairs with a model that also sets upside_down and no rotation happens.

resample

attribute[source]
resample: Resample

Resize algorithm the model’s training pipeline used. An un-suffixed name has OpenCV/torch semantics, an _aa suffix has PIL’s (an antialiased filter whose support widens with the downscale factor): "bilinear" (4-tap half-pixel-center bilinear; the default, which most trained policies match), "bilinear_aa" / "bicubic_aa" / "lanczos3_aa" (PIL’s BILINEAR / BICUBIC / LANCZOS), or "area" (OpenCV INTER_AREA). Bare "bicubic"/"lanczos3" are not accepted – name the library whose kernel you trained against.

allow_upscale

attribute[source]
allow_upscale: bool

Permit a target larger than the env’s native resolution (interpolating detail that is not there). Off by default: an upscaling target is a resolve error unless this is set.

fit

attribute[source]
fit: FitMode | Sequence[FitMode] | None

How to reconcile a target whose aspect ratio differs from the env image: "stretch" (distort), "crop" (cover + center-crop), or "pad" (letterbox) – or a sequence of them in preference order. The resolver picks, per env, the first that does not need a disallowed upscale, so one spec can crop a large camera and letterbox a small one. Required only on an aspect mismatch; absent it, an aspect-changing resize is a resolve error.

optional

attribute[source]
optional: bool

Zero-fill a black frame when the env does not provide this camera, instead of failing resolution. Needs height, width, and channels so the blank can be sized.

fill

attribute[source]
fill: int | None

Fill value (0-255) for the blank frame an optional camera synthesizes when the env offers no matching image; default 0 (black). Requires optional=True (without it the fill could never take effect, so setting it alone is a construction error); named to match Actuator.fill.

stack

attribute[source]
stack: int

Number of frames to stack on a new leading axis (frame history). 1 (default) means no stacking. Stacking is applied by the adapter core from an episode-keyed rolling window, padded at the start of an episode (see stack_pad) and cleared on reset. Which frames it gathers is offsets (or stride); by default they are the last stack consecutive ones. The env still sends one frame per step either way.

crop

attribute[source]
crop: float | None

Side fraction of the frame a center crop keeps, in (0, 1] (2/3 keeps the middle two thirds of each axis). Pass crop or crop_area, not both.

crop_area

attribute[source]
crop_area: float | None

The same center crop stated as an area fraction, in (0, 1] – the form training pipelines usually quote (“a 90% center crop”). The side fraction is its square root (0.9 -> 0.949).

crop_mode

attribute[source]
crop_mode: CropMode

How the crop box is taken. "zoom" (the default) resamples the fractional box straight to the target – PIL’s Image.resize(size, box=...), one pass, no intermediate rounding. "slice" cuts an integer center box out first and resizes that, which is what a numpy slice in a training pipeline does.

jpeg_quality

attribute[source]
jpeg_quality: int | None

Quality of a JPEG round-trip applied to the upright frame before the crop and resize, on the IJG 1-100 scale (95 is the usual training value). A declaration, not a request: a pipeline that stored its frames as JPEG fed the model the codec’s artifacts, so the adapter reproduces them rather than handing the model a cleaner frame than it was trained on. Baseline sequential, 4:2:0 box-averaged chroma, standard IJG tables; needs a 3-channel image.

channel_order

attribute[source]
channel_order: ChannelOrder

Channel order the model was trained on. "bgr" swaps red and blue after the spatial ops and before the dtype cast, and needs a 3-channel image.

render

attribute[source]
render: int | tuple[int, int] | None

The camera resolution this model was trained against, as a square int or an (height, width) pair (each axis 1-4096). An assertion, not a request: the adapter never resizes a camera to reach it. The platform binds the environment’s camera dial from it before the run; resolution then fails with RenderMismatch if the camera it bound does not actually render at this size. Declare it instead of pinning the env’s camera width/height by hand.

offsets

attribute[source]
offsets: StackSpec | None

The frame window stack gathers, as non-positive offsets from the current step – oldest first, ending at 0. (-6, -4, -2, 0) is “every second frame of the last seven”. None (the default) is the contiguous window stack already describes. offsets and stack are never inferred from each other: both are declared and they must agree (len(offsets) == stack), and the window may span at most 128 consecutive frames.

stack_pad

attribute[source]
stack_pad: StackPad

What fills the window before the episode has produced enough frames: "first" (the default) replicates the first observed frame, "black" a raw 8-bit 0 frame pushed through this input’s pipeline – so under normalize=(-1.0, 1.0) a black pad frame is -1.0, not 0.0. Requires stack > 1.

part

attribute[source]
part: str | None

The body part the wanted camera sits on, when the role repeats across a body ("head", "left_arm", …): an identity key the resolver matches on. Naming one binds only that camera; naming none binds the env’s only camera of the role under any part (with an info) and fails when there are several. Keyword-only and omitted from the wire when unset.