Qwen2.5-3B-Instruct vs Qwen-72B

Try Try

Qwen2.5-3B-Instruct has 3.09B parameters (3.09B active per token) and takes 6.17 GB on disk; Qwen-72B has 72.3B (71.0B active) and takes 145 GB. Compare them layer by layer.

At a glance

Qwen/Qwen2.5-3B-InstructQwen/Qwen-72B
Repo createdSep 17, 2024updated Sep 25, 2024Nov 26, 2023updated Oct 9, 2024
Model typeDenseFull attention (GQA)DenseFull attention (MHA)
InputsTextText
Total parameters3.09B23× less72.3B23× more
Active per token3.09B23× less100% of the model71.0B23× more98% of the model
ExpertsNone (dense)None (dense)
Max context (from config)32K tokenssame32K tokenssame
Layers362.2× less802.2× more
On disk6.17 GB23× less2 files145 GB23× more82 files
PrecisionBF16 (100%)BF16 (100%)
QuantizationNoneoriginal precision (BF16)Noneoriginal precision (BF16)
Fewest GPUs1× RTX 4060a single 8 GB card · weights only1× NVIDIA B200a single 180 GB card · weights only
Licenseqwen-research (custom)tongyi-qianwen-license-agreement (custom)
GitHubQwenLM/Qwen2.5QwenLM/Qwen