Instructions to use xpuenabler/MolmoAct2-LIBERO-LP-VTP-SD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use xpuenabler/MolmoAct2-LIBERO-LP-VTP-SD with LeRobot:
- Notebooks
- Google Colab
- Kaggle
MolmoAct2-LIBERO-LP-VTP-SD โ ํตํฉ ์์ถ (Layer Prune + Vision Token Prune + Step Decrease)
allenai/MolmoAct2-LIBERO (5.44B VLA)์ ์ธ ๊ฐ์ง ์์ถ ๊ธฐ๋ฒ์ ๋์ ์ ์ฉํด LIBERO๋ก fine-tuneํ LeRobot ํฌ๋งท ์ฒดํฌํฌ์ธํธ.
| ๊ธฐ๋ฒ | ๋ด์ฉ | ํจ๊ณผ |
|---|---|---|
| LP (layer pruning) | ViT 25L โ 10L ํ์(xpuenabler/MolmoAct2-LIBERO-ViT10L)์์ ์์ | ViT ํ๋ผ๋ฏธํฐ 439M โ 212M |
| VTP (vision token pruning) | Grid Sampler, ์ด๋ฏธ์ง๋น K=16 ํ ํฐ | prefill seq_len 489 โ 129 |
| SD (step decrease) | flow-matching num_flow_timesteps=2 ํ์ต, 2-step ๋์ฝ๋ / n_action_steps=2 |
action head ๋์ฝ๋ ~5ร ์ถ์ |
ํ์ต
- ์ฝ๋:
nota-github/xpu-lerobot@feat/lp+vtp+sd-libero(scripts/train_molmoact2_libero_vtp+lp+sd.sh) - ๋ฐ์ดํฐ:
allenai/MolmoAct2-LIBERO-Dataset, 20000 steps, effective batch 32 (4รB200 DDP, per-GPU 8), bf16, seed 1000 - Grid Sampler ๊ฐ์ค์น๋ random init์์ ํ์ต, ๋๋จธ์ง๋ ViT10L ํ์์์ warm-start
์ฑ๋ฅ
- LIBERO 4-suite ํ๊ฐ ๋ฐ module-wise latency๋ Confluence ๋ฌธ์ "Pruning/Distilation ์ข ํฉ (LP+VTP+SD) MolmoAct2 ์ ์ฉ ๋ฐ ์ฑ๋ฅ ํ๊ฐ" ์ฐธ๊ณ (10000-step ์ฒดํฌํฌ์ธํธ ๊ธฐ์ค ์ธก์ ; ๋ณธ repo๋ ์ต์ข 20000-step).
- B200 batch=1 ์ฐธ๊ณ ์น (2-step ๋์ฝ๋): TOTAL ~94ms/decision (vision 5.8 / llm 20.3 / action head 52.4 ms).
์ฌ์ฉ (LeRobot)
lerobot-eval \
--policy.path=xpuenabler/MolmoAct2-LIBERO-LP-VTP-SD \
--policy.inference_action_mode=continuous \
--policy.num_inference_steps=2 \
--env.type=libero ...
์ฃผ์: Grid Sampler๋ single-crop(crop_mode="resize") ์ ์ฉ.
- Downloads last month
- 50