N0-TWAM β€” A Tactile-Native World Action Model

$N_0$-TWAM is a Vision–Tactile–Action world-action model. Vision, tactile, and action are jointly modeled by a Mixture-of-Transformers (MoT) under a single rectified-flow / flow-matching objective: the model predicts the visual and tactile future and generates the low-level action that realizes it.

This repository holds the pretrained weights as a ready-to-use bundle. The training, inference-server, and post-training code lives at πŸ‘‰ https://github.com/neoteai/N0-TWAM

Contents

A self-contained bundle β€” everything the model needs to load and serve:

.
β”œβ”€β”€ transformer/              # the trained N0-TWAM weights (MoT, 20-d action, tactile)
β”œβ”€β”€ vae/                      # Wan2.2 AutoencoderKLWan (48-channel)
β”œβ”€β”€ text_encoder/             # umT5-xxl (frozen)
β”œβ”€β”€ tokenizer/
β”œβ”€β”€ norm_stat_pretrain.json   # action q01/q99 the checkpoint was trained with β€”
β”‚                             #   the server de-normalizes actions with these
β”‚                             #   (auto-loaded from the bundle root by twam_server)
└── empty_emb.pt              # empty-prompt text embedding (used by the
                              #   post-training dataloader's CFG text drop)

Usage

1. Install the code

git clone https://github.com/neoteai/N0-TWAM.git
cd N0-TWAM
pip install .
pip install flash-attn --no-build-isolation

2. Download this bundle

from huggingface_hub import snapshot_download
bundle = snapshot_download("NeoteAI/n0-twam-base")
# `bundle` now contains transformer/ vae/ text_encoder/ tokenizer/
# plus norm_stat_pretrain.json and empty_emb.pt

3a. Load the model

import torch
from n0_twam.models.utils import load_mot_checkpoint

model = load_mot_checkpoint(f"{bundle}/transformer",
                            torch_dtype=torch.bfloat16, torch_device="cuda")
print(f"{sum(p.numel() for p in model.parameters()) / 1e9:.2f} B params")
# -> 7.16 B params

3b. Or serve it (observation β†’ action)

Point the inference-server config at this downloaded bundle β€” it already contains every component the server needs, including the training-time action norm stats (norm_stat_pretrain.json, auto-loaded from the bundle root) β€” then launch the websocket server:

# in n0_twam/configs/twam_server_cfg.py
twam_server_cfg.wan22_pretrained_model_name_or_path = "<path to the downloaded bundle>"
export PYTHONPATH=$PWD:$PWD/n0_twam CUDA_VISIBLE_DEVICES=0
export RANK=0 LOCAL_RANK=0 WORLD_SIZE=1 MASTER_ADDR=127.0.0.1 MASTER_PORT=29988
python -m n0_twam.n0_twam_server --config-name twam_server --port 29601

Query it from a client (a reset with a language prompt, then per-step observations):

import numpy as np
from n0_twam.utils.Simple_Remote_Infer.deploy.websocket_client_policy import WebsocketClientPolicy

client = WebsocketClientPolicy("127.0.0.1", 29601)
client.infer({"reset": True, "prompt": "pick up the object"})

cams = ["observation.images.third_view",
        "observation.images.left_wrist_view",
        "observation.images.right_wrist_view"]
frame = {k: np.zeros((256, 256, 3), np.uint8) for k in cams}   # your RGB frames
state = np.zeros(20, np.float32)                                # your current EE state
action = client.infer({"obs": [frame], "current_state": state})["action"]  # (20, 2, 16)

See DEPLOY.md for the full serving guide, and POST_TRAINING.md to fine-tune the model on your own robot.

Model details

Architecture 3-expert Mixture-of-Transformers (video / tactile / action) with shared attention
Parameters ~7.16 B (bf16)
Backbone WAN2.2 TI2V-5B video diffusion transformer
Video VAE Wan2.2 AutoencoderKLWan (z_dim=48, 4Γ— temporal / 16Γ— spatial)
Text encoder umT5-xxl (4096-d), frozen
Objective Rectified-flow / flow-matching, per-frame timesteps
Action space 20-dim dual-arm end-effector, Ο€0.5-style horizon delta
Tactile global (co-generated diffusion target) + optional local (observed input)

License

Released under the CC-BY-NC-SA-4.0 license (see LICENSE). The vae/, text_encoder/ and tokenizer/ components are redistributed from Wan2.2 and keep their original Apache 2.0 license and notices.

Acknowledgments

Builds upon LingBot-VA, Wan2.2, FastWAM (MoT design), and LeRobot.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading