rlmesh.adapters.Image
An image input expected by a model.
Image(
role: str,
height: int | None = None,
width: int | None = None,
layout: ImageLayout = 'hwc',
channels: int | None = None,
dtype: str = 'uint8',
normalize: bool | tuple[float, float] = False,
lead_dims: int = 0,
upside_down: bool = False,
resample: Resample = 'bilinear',
allow_upscale: bool = False,
fit: FitMode | Sequence[FitMode] | None = None,
optional: bool = False,
fill: int | None = None,
stack: int = 1,
size: InitVar[int | None] = None,
*,
crop: float | None = None,
crop_area: float | None = None,
crop_mode: CropMode = 'zoom',
jpeg_quality: int | None = None,
channel_order: ChannelOrder = 'rgb',
render: int | tuple[int, int] | None = None,
offsets: StackSpec | None = None,
stack_pad: StackPad = 'first',
part: str | None = None,
stride: InitVar[int | None] = None,
)There is no key – placement in the input tree is the payload position.
Attributes
- roleSemantic role matched against env image features.
- heightTarget image height, or None to keep the env height.
- widthTarget image width, or None to keep the env width.
- layoutAxis layout the model expects (
"hwc"the default, or"chw"). - channelsChannel count the model expects (e.g. 3 for RGB, 1 for grayscale).
- dtypeNumPy dtype name the model expects.
- normalizeWhether (and into what range) to map 8-bit pixel values before casting:
False(off, the default),True(the conventional[0, 1]), or a(low, high)pair for a model trained on a different range (e.g. - lead_dimsNumber of leading singleton axes to add (batch/time).
- upside_downWhether the model was trained on images rotated 180 degrees from the canonical upright orientation (a true rotation -- rows and columns reversed -- not a vertical flip).
- resampleResize algorithm the model's training pipeline used.
- allow_upscalePermit a target larger than the env's native resolution (interpolating detail that is not there).
- fitHow to reconcile a target whose aspect ratio differs from the env image:
"stretch"(distort),"crop"(cover + center-crop), or"pad"(letterbox) -- or a sequence of them in preference order. - optionalZero-fill a black frame when the env does not provide this camera, instead of failing resolution.
- fillFill value (0-255) for the blank frame an
optionalcamera synthesizes when the env offers no matching image; default 0 (black). - stackNumber of frames to stack on a new leading axis (frame history).
- cropSide fraction of the frame a center crop keeps, in
(0, 1](2/3keeps the middle two thirds of each axis). - crop_areaThe same center crop stated as an *area* fraction, in
(0, 1]-- the form training pipelines usually quote ("a 90% center crop"). - crop_modeHow the crop box is taken.
- jpeg_qualityQuality of a JPEG round-trip applied to the upright frame *before* the crop and resize, on the IJG 1-100 scale (
95is the usual training value). - channel_orderChannel order the model was trained on.
- renderThe camera resolution this model was trained against, as a square
intor an(height, width)pair (each axis 1-4096). - offsetsThe frame window
stackgathers, as non-positive offsets from the current step -- oldest first, ending at0. - stack_padWhat fills the window before the episode has produced enough frames:
"first"(the default) replicates the first observed frame,"black"a raw 8-bit0frame pushed through this input's pipeline -- so undernormalize=(-1.0, 1.0)a black pad frame is-1.0, not0.0. - partThe body part the wanted camera sits on, when the role repeats across a body (
"head","left_arm", ...): an identity key the resolver matches on.
role
attribute[source]role: strSemantic role matched against env image features.
height
attribute[source]height: int | NoneTarget image height, or None to keep the env height.
width
attribute[source]width: int | NoneTarget image width, or None to keep the env width.
layout
attribute[source]layout: ImageLayoutAxis layout the model expects ("hwc" the default, or "chw"). This is the model author’s declaration: when it differs from the env’s layout the adapter transposes silently to match, so a model that wants "chw" must say so – an omitted layout means "hwc" and the env’s frame is fed through unchanged. (channels guards the channel count, not the axis order.)
channels
attribute[source]channels: int | NoneChannel count the model expects (e.g. 3 for RGB, 1 for grayscale). When set, a resolve error if the env image differs; the adapter does not convert between channel counts.
dtype
attribute[source]dtype: strNumPy dtype name the model expects.
normalize
attribute[source]normalize: bool | tuple[float, float]Whether (and into what range) to map 8-bit pixel values before casting: False (off, the default), True (the conventional [0, 1]), or a (low, high) pair for a model trained on a different range (e.g. (-1.0, 1.0)). One field, so an on/off flag can never disagree with a range; False is an authoritative off-switch.
lead_dims
attribute[source]lead_dims: intNumber of leading singleton axes to add (batch/time).
upside_down
attribute[source]upside_down: boolWhether the model was trained on images rotated 180 degrees from the canonical upright orientation (a true rotation – rows and columns reversed – not a vertical flip). Declared on both ends: the adapter rotates only when the env and model disagree, so an env that already renders upside-down pairs with a model that also sets upside_down and no rotation happens.
resample
attribute[source]resample: ResampleResize algorithm the model’s training pipeline used. An un-suffixed name has OpenCV/torch semantics, an _aa suffix has PIL’s (an antialiased filter whose support widens with the downscale factor): "bilinear" (4-tap half-pixel-center bilinear; the default, which most trained policies match), "bilinear_aa" / "bicubic_aa" / "lanczos3_aa" (PIL’s BILINEAR / BICUBIC / LANCZOS), or "area" (OpenCV INTER_AREA). Bare "bicubic"/"lanczos3" are not accepted – name the library whose kernel you trained against.
allow_upscale
attribute[source]allow_upscale: boolPermit a target larger than the env’s native resolution (interpolating detail that is not there). Off by default: an upscaling target is a resolve error unless this is set.
fit
attribute[source]fit: FitMode | Sequence[FitMode] | NoneHow to reconcile a target whose aspect ratio differs from the env image: "stretch" (distort), "crop" (cover + center-crop), or "pad" (letterbox) – or a sequence of them in preference order. The resolver picks, per env, the first that does not need a disallowed upscale, so one spec can crop a large camera and letterbox a small one. Required only on an aspect mismatch; absent it, an aspect-changing resize is a resolve error.
optional
attribute[source]optional: boolZero-fill a black frame when the env does not provide this camera, instead of failing resolution. Needs height, width, and channels so the blank can be sized.
fill
attribute[source]fill: int | NoneFill value (0-255) for the blank frame an optional camera synthesizes when the env offers no matching image; default 0 (black). Requires optional=True (without it the fill could never take effect, so setting it alone is a construction error); named to match Actuator.fill.
stack
attribute[source]stack: intNumber of frames to stack on a new leading axis (frame history). 1 (default) means no stacking. Stacking is applied by the adapter core from an episode-keyed rolling window, padded at the start of an episode (see stack_pad) and cleared on reset. Which frames it gathers is offsets (or stride); by default they are the last stack consecutive ones. The env still sends one frame per step either way.
crop
attribute[source]crop: float | NoneSide fraction of the frame a center crop keeps, in (0, 1] (2/3 keeps the middle two thirds of each axis). Pass crop or crop_area, not both.
crop_area
attribute[source]crop_area: float | NoneThe same center crop stated as an area fraction, in (0, 1] – the form training pipelines usually quote (“a 90% center crop”). The side fraction is its square root (0.9 -> 0.949).
crop_mode
attribute[source]crop_mode: CropModeHow the crop box is taken. "zoom" (the default) resamples the fractional box straight to the target – PIL’s Image.resize(size, box=...), one pass, no intermediate rounding. "slice" cuts an integer center box out first and resizes that, which is what a numpy slice in a training pipeline does.
jpeg_quality
attribute[source]jpeg_quality: int | NoneQuality of a JPEG round-trip applied to the upright frame before the crop and resize, on the IJG 1-100 scale (95 is the usual training value). A declaration, not a request: a pipeline that stored its frames as JPEG fed the model the codec’s artifacts, so the adapter reproduces them rather than handing the model a cleaner frame than it was trained on. Baseline sequential, 4:2:0 box-averaged chroma, standard IJG tables; needs a 3-channel image.
channel_order
attribute[source]channel_order: ChannelOrderChannel order the model was trained on. "bgr" swaps red and blue after the spatial ops and before the dtype cast, and needs a 3-channel image.
render
attribute[source]render: int | tuple[int, int] | NoneThe camera resolution this model was trained against, as a square int or an (height, width) pair (each axis 1-4096). An assertion, not a request: the adapter never resizes a camera to reach it. The platform binds the environment’s camera dial from it before the run; resolution then fails with RenderMismatch if the camera it bound does not actually render at this size. Declare it instead of pinning the env’s camera width/height by hand.
offsets
attribute[source]offsets: StackSpec | NoneThe frame window stack gathers, as non-positive offsets from the current step – oldest first, ending at 0. (-6, -4, -2, 0) is “every second frame of the last seven”. None (the default) is the contiguous window stack already describes. offsets and stack are never inferred from each other: both are declared and they must agree (len(offsets) == stack), and the window may span at most 128 consecutive frames.
stack_pad
attribute[source]stack_pad: StackPadWhat fills the window before the episode has produced enough frames: "first" (the default) replicates the first observed frame, "black" a raw 8-bit 0 frame pushed through this input’s pipeline – so under normalize=(-1.0, 1.0) a black pad frame is -1.0, not 0.0. Requires stack > 1.
part
attribute[source]part: str | NoneThe body part the wanted camera sits on, when the role repeats across a body ("head", "left_arm", …): an identity key the resolver matches on. Naming one binds only that camera; naming none binds the env’s only camera of the role under any part (with an info) and fails when there are several. Keyword-only and omitted from the wire when unset.