SmolVLM2-500M-Video-Instruct

Try Try

SmolVLM2-500M-Video-Instruct has 507M parameters (374M active per token) and takes 2.03 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

HuggingFaceTB/SmolVLM2-500M-Video-Instruct
Repo createdFeb 11, 2025updated Apr 8, 2025
Model typeDenseFull attention (GQA)
InputsText + images86.4M vision encoder
Total parameters507M
Active per token374M74% of the model
ExpertsNone (dense)
Max context (from config)8K tokens
Layers32
On disk2.03 GB1 file
PrecisionFP32 (100%)
QuantizationNoneoriginal precision (FP32)
Fewest GPUs1× RTX 4060a single 8 GB card · weights only
Licenseapache-2.0
GitHubNot linked