mm
Video understanding, in a much smaller package
Hugging Face’s SmolVLM2 release explored video understanding with models in the 2.2B, 500M, and 256M size range.
That makes a useful counterpoint to the biggest-model headlines: deployment constraints can be part of the research target, rather than something considered after training.
DEMO · SOURCED SUMMARYHugging Face