↗ toy world/ lab
Connecting to the labSIMULATION LAB
AN INTERACTIVE EXPERIMENT IN LEARNED DYNAMICS

A small world.
A model that learns it.

Move a pointer. Record what happens. Teach a model to imagine the next move—and use those predictions to act.

01 / THE PROGRAMMED WORLD

Get a feel for the physics.

No AI in this step
THE CONCEPT · DYNAMICS

A world changes over time.

Dynamics describe how a world’s current state and an action produce its next state. Actions are only part of the story: motion can continue even when an agent does nothing.

IN THIS EXPERIMENT

Angle, speed, and a small push.

The pointer’s state is its angle and angular speed. Left and right apply drive; hold applies none. Damping slows it down. The simulator uses programmed equations to produce the “ground truth” we will later learn from.

Current angle + speed+left / hold / right→next angle + speed
Try this—and understand what you are seeing

Start the simulation, choose Right, then Hold before reaching the travel limit. The pointer coasts because it already has speed. Increase damping and repeat: it stops sooner. A single picture tells you its position, but not whether it is moving left or right.

State versus observation: the simulator internally knows the exact angle and speed. A camera only provides an image—an observation. Later steps estimate motion from consecutive images instead of handing the model that hidden state. This first sandbox contains no learning.

LIVE SIMULATION0.0 s
↗ Blue pointer Red target10 steps / second
Next: turn this experience into training examples.
02 / THE DATASET

Every action leaves an example.

Loading dataset
THE CONCEPT · EXPERIENCE

Record causes alongside consequences.

A transition records what was observed, which action was taken, and what was observed next. An episode is a sequence of these transitions from one starting condition. Action labels help the model distinguish how different inputs affect motion.

IN THIS EXPERIMENT

Random actions become training examples.

We render images, locate the blue pointer, and estimate speed from its change in angle. Each saved example contains those measurements, a left/hold/right action, and the next measurements. Each episode varies the starting angle and motor behavior.

Image → measured angle & speed+recorded action→next measured state
Why random actions, and how is this different from imitation?

Dynamics learning predicts the consequence of an action. Random exploration is useful because we need examples of many movements, not just successful ones. Browse transitions: the same action can have different effects depending on the pointer’s existing speed.

Imitation learning instead predicts the action a demonstrator would choose in a situation. We are not doing that here: there is no expert player to copy and no training objective to pick winning moves. The data teaches how the world responds, not which goal to pursue.

Coverage matters: behavior never seen during collection may be predicted poorly. We vary motor strength and damping to expose the model to more than one exact simulated setup. Saved features are measurements from pixels; raw images are not retained in this dataset.

A real transition from your dataset

Observed state
→ACTIONRight
Next observed state
States are measured from rendered pixels.
Next: learn the relationship between action and consequence.
03 / LEARNING

From examples to predictions.

Loading checkpoint
THE CONCEPT · SUPERVISED LEARNING

Make a prediction. Measure the mistake.

A neural network starts with untrained parameters. For each example, it predicts the next state. A loss measures the difference from the recorded outcome; training adjusts the parameters to reduce that loss. An epoch is one pass through the training examples.

IN THIS EXPERIMENT

A tiny network learns pointer motion.

The inputs are measured angle, speed, and action. Two layers of 64 units predict the changes in angle and speed. The network is not given the motion equations. We check predictions on held-out episodes and save the checkpoint with the lowest validation loss.

Predict next state→compare with recorded next state→adjust network parameters
Read the learning curve and understand the limits

Training loss measures error on examples used to update the model. Validation loss measures error on episodes it does not train on. If training error falls while validation error rises, the model may be fitting the training examples without improving on new ones. Holding out entire episodes avoids mixing neighboring frames from the same recording into both sets.

The chart combines angle and speed errors after scaling them. The degree metric is easier to interpret: it measures next-angle error against the image-derived measurement. It does not measure success at reaching a target, and it does not include all camera measurement error.

What is learned? Only the dynamics. Color detection and image rendering are still programmed. MIRA learns to generate visual outcomes; this smaller, structured model lets us study the prediction-and-control loop before adding learned vision. Motor strength and damping are hidden inputs, so this model learns a typical response across their training range.

INPUTAngle · speed · action
→
SMALL NEURAL NETWORK64 units → 64 units
→
OUTPUTNext angle · next speed

Learning curve

Training Validation

Normalized prediction error · lower is better · entire episodes are held out for validation

BEST VALIDATION ANGLE ERROR—
EPOCHS COMPLETED—
TRAINING LOOP—
Next: see whether predictions hold up beyond one step.
04 / OPEN-LOOP PREDICTION

Can it imagine what happens next?

No correction after the first frame
THE CONCEPT · A WORLD MODEL

Ask “what happens if…?”

A world model predicts how an environment evolves under actions. In an open-loop rollout, each predicted state becomes the input to the next prediction. There are no fresh observations to correct mistakes, so even small errors can accumulate.

IN THIS EXPERIMENT

Two worlds receive the same commands.

The left dial comes from the programmed simulator. The right dial comes from the trained network’s predictions, drawn by a fixed renderer. Both start from the same measured state and receive 20 actions. Only the simulator continues using the true motion equations.

Initial observation→predicted state 1→predicted state 2→…
What to look for, and why learn a simulator we already have?

Compare Drive right → coast with Drive right → brake left → hold. Scrub the timeline and watch when the two pointers diverge. The error shown averages their angle differences over the rollout. It is not the validation error from the previous step: predicting many steps without correction is a harder test.

This model is not choosing the actions. You provide them. It predicts consequences, without deciding which consequence is desirable. That decision is the next step.

For this tiny task, the programmed simulator is already cheap and accurate. The learned copy is useful as a learning experiment, not because it is a better replacement. The wider motivation is learning usable dynamics from experience when writing an accurate simulator is difficult. Success here does not prove that broader capability or physical transfer.

ACTUAL SIMULATOR
LEARNED PREDICTION
0.0 s
MEAN PREDICTION ERROR—
HORIZON2.0 s
Next: use imagined futures to choose actions.
05 / MODEL PREDICTIVE CONTROL

Imagine. Choose. Observe again.

256 candidate sequences per decision
THE CONCEPT · PLANNING WITH A MODEL

Predictions become useful decisions.

A world model predicts outcomes; a controller chooses actions to reach a goal. Model predictive control imagines several action sequences, scores their consequences, executes only the first action of the best sequence, then observes again and replans.

IN THIS EXPERIMENT

Choose the move that approaches red.

The planner uses the learned network to imagine 256 sequences, each eight steps long. It penalizes distance from the target, excessive motion, and predicted travel-limit violations. After one action in the actual simulator, a new image-derived state replaces the prediction.

Observe→imagine & score futures→act once→observe again ↻
How to judge success—and what changes for a real robot

No separate policy network is trained here. The action chooser is a programmed planner using learned dynamics. A different design could train a policy to choose actions directly. The model and the mechanism that selects actions are distinct components.

Change motor strength or damping beyond the training ranges to test distribution shift: the controller now faces unfamiliar dynamics. Fresh observations help it recover from prediction errors, but do not guarantee success. Compare with the simple visual controller below; a learned model is not automatically better.

Success is defined narrowly: within about 4.6° of the target for one continuous second at any point in a ten-second trial. A pointer can satisfy that and drift later, which is why final error is also shown. The playback shows a computed simulation run, not a live robot.

Sim-to-real comes next: replace rendered observations with camera measurements and map actions to a servo. Camera geometry, latency, and motor response must match closely enough—or require adaptation. Physical transfer is not implemented or verified by these simulation results.

Let the learned model guide the pointer

Ready to run
0.0 s
TARGET HOLD—
FINAL ERROR—
CHOSEN ACTION—
THE BENCHMARK

Does the learned controller actually help?

LEARNED-MODEL CONTROLLER—

SIMPLE VISUAL CONTROLLER—

OPEN-LOOP PREDICTION ERROR—

Across two-second imagined rollouts

Success: within 4.6° for one continuous second during a 10-second trial. These are simulation results, not hardware results.

Next experiment: camera + cardboard pointer + a real servo.Physical transfer has not been tested.