Q-50M-Base
A 50.9M-parameter decoder-only language model, pretrained from scratch on 5 billion tokens of FineWeb-Edu. A base model for continuation, study, and fine-tuning — small enough to run anywhere.
- Parameters
- 50.9M
- Training tokens
- 5B
- Context
- 2,048
- Vocab
- 32,768 BPE
- Attention
- GQA · QK-Norm
- Val. perplexity
- 24.47