Back to the stream
mm

Video understanding, in a much smaller package

Hugging Face’s SmolVLM2 release explored video understanding with models in the 2.2B, 500M, and 256M size range.

That makes a useful counterpoint to the biggest-model headlines: deployment constraints can be part of the research target, rather than something considered after training.

Read the technical details and examples →

DEMO · SOURCED SUMMARYHugging Face