tiny-Qwen2ForCausalLM-2.5 vs Qwen3-8B

Try Try

tiny-Qwen2ForCausalLM-2.5 has 2.44M parameters (1.22M active per token) and takes 4.87 MB on disk; Qwen3-8B has 8.19B (7.57B active) and takes 16.4 GB. Compare them layer by layer.

At a glance

trl-internal-testing/tiny-Qwen2ForCausalLM-2.5Qwen/Qwen3-8B
Repo createdNov 25, 2024updated Sep 8, 2026Apr 27, 2025updated Jul 26, 2025
Model typeDenseFull attention (GQA)DenseFull attention (GQA)
InputsTextText
Total parameters2.44M3364× less8.19B3364× more
Active per token1.22M6211× less50% of the model7.57B6211× more92% of the model
ExpertsNone (dense)None (dense)
Max context (from config)32K tokens1.2× less40K tokens1.2× more
Layers218× less3618× more
On disk4.87 MB3364× less1 file16.4 GB3364× more5 files
PrecisionBF16 (100%)BF16 (100%)
QuantizationNoneoriginal precision (BF16)Noneoriginal precision (BF16)
Fewest GPUs1× RTX 4060a single 8 GB card · weights only1× RTX 4090a single 24 GB card · weights only
LicenseNot statedapache-2.0
GitHubNot linkedQwenLM/Qwen3