Superhuman/X ArchiveView on X
Generalist

@GeneralistAI

Introducing GEN-1.5, a one-shot learner.

It can learn new tasks in a few seconds. Show it what to do, and it generalizes.

This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
2741.5K10.4K5.1K
Generalist

@GeneralistAI

GEN-1.5, our latest embodied foundation model, can learn new tasks prompted with 3 - 12 seconds of a single demonstration, no gradient updates or fine-tuning. It generalizes prompts to new situations, recovers from mistakes, and improvises new strategies to reach the same goal.
633688117
Generalist

@GeneralistAI

Physical prompts can be composed. Demonstrations of 2 different tasks in context prompts GEN-1.5 to chain them into one continuous skill. The model bridges them and produces intermediate motions (repositioning, regrasping, error recovery) that appear in neither demonstration.
1831134
Generalist

@GeneralistAI

In-context learning also crosses the sim-to-real gap, zero-shot. Prompts can be formed entirely from simulated experience (e.g., from a scripted policy, an RL agent, or a human teleoperating a simulated robot) and be used to produce behaviors on a real robot. The model was not trained on the task in either the simulator or the real world.
2827645
Generalist

@GeneralistAI

In some cases, in-context learning with GEN-1.5 transfers across the embodiment gap entirely: a human demonstrates a task with their own hands, observable through the robot’s cameras, and the robot can reproduce it immediately afterward.
51432843
Generalist

@GeneralistAI

For few-shot learning, it can adapt to new physical tasks in as few as 1 - 10 gradient steps on 1 - 5 minutes of data (~10 - 50 demonstrations). In practice, this can be described as test-time training in a low-data regime. We did not tune this procedure or sweep hyperparameters; these results come largely out of the box.
Image from the post
112331113
Generalist

@GeneralistAI

Experiments across 10 diverse tasks show 59% average success with one-shot physical prompting, straight from pretraining. With few-shot learning, performance rises to 83% via 10 gradient steps on 5 minutes of data per task.

Although the tasks are simple and success rates are modest, it’s the first model we know of that exhibits the general ability to learn a wide range of dexterous closed-loop physical tasks from just one or few demonstrations. This accelerates reaching a base level of competence for new skills that can be subsequently refined towards mastery.
1418618
Generalist

@GeneralistAI

Fine-tuned (or prompted) behaviors generalize beyond their demonstrations, and can improvise fundamentally different manipulation strategies to achieve the same goal.

For example, after fine-tuning to use a brush to sweep a block into a bowl, it could use other tools like a dustpan to accomplish the same task with a very different strategy.
1215814
Generalist

@GeneralistAI

Or when fine-tuned to place a block into a bowl, it can clear obstacles (like a piece of paper covering the bowl) to complete the task, despite that not being in the demonstrations.
1215716
Generalist

@GeneralistAI

When a Lego brick gets unexpectedly stuck on the fingertips, the model uses the other hand to remove them.
2720414
Generalist

@GeneralistAI

The model sometimes uses both hands to rotate a jar lid, with a fundamentally different contact and motion strategy than the fine-tuning demonstrations.
1113413
Generalist

@GeneralistAI

Here’s an uncut video of prompting the model to perform 2 different tasks back-to-back: (i) unzipping a pencil pouch, and (ii) retrieving money from the pouch.
1313616
Generalist

@GeneralistAI

GEN-1.5 has been training continuously for over 8 months. We left it running because every metric we tracked kept improving with the engine: absorbing more data, scaling more efficiently, boosting post-training, and compounding step-change improvements through algorithmic advances.
Image from the post
41429455
Generalist

@GeneralistAI

To us, GEN-1.5 represents a new frontier of generality — one that challenges our own understanding of how these models behave when pretrained at a scale of physical interaction data few thought possible without shortcuts. We do not yet see where this asymptotes.

Read more in the full blog: generalistai.com/blog/gen-1.5
41928199
End of thread