TacFiLM overview: tactile, visual, and language inputs, baseline fusion approaches, and the TacFiLM-augmented VLA.

Tactile Modality Fusion for Vision-Language-Action Models

TacFiLM is a lightweight post-training fusion method that conditions a VLA’s intermediate visual features on pretrained tactile representations through feature-wise linear modulation, improving success rate, completion time, and force stability on contact-rich insertion and drawer-opening tasks.

March 2026 · Anas Houssaini