How CoLT Accelerates Multimodal Reasoning Over 20× with a 3‑Step Latent Thought Chain
CoLT replaces the verbose textual chain‑of‑thought with just three latent vectors, delivering up to 22.6× faster generation and 10.1× end‑to‑end inference speedups while achieving a 79.1% average accuracy across eight multimodal benchmarks, all without any auxiliary visual annotations.
