Small language models
Small-model research includes sub-billion-parameter language models for mobile deployment.
MobileLLM studies sub-billion-parameter models for mobile deployment. TinyStories trains models below ten million parameters.
Phi-1 has 1.3 billion parameters and uses selected textbook-quality data. DistilBERT uses knowledge distillation and reduced BERT model size by 40 percent while retaining 97 percent of its reported language-understanding capability.
MobileLLM emphasizes model architecture at the sub-billion scale. The Gemma paper includes 2 billion and 7 billion parameter models.
Large pretrained models can be difficult to run under constrained compute budgets.
Compare parameter scale with the deployment budget.
- 1 · boundSet the deployment resource envelope.
- 2 · sizeChoose a parameter count that fits that envelope.
- 3 · compareModel architecture matters at sub-billion scale.
- 4 · measureRun a comparative on-device study.
TinyStories studies models below ten million parameters; phi-1 has 1.3 billion parameters.
| Who | What they ask | What it works with |
|---|---|---|
| Mobile developer | “Can this model run within the device budget?” | Parameters, memory and measured device latency |
| Model trainer | “Should the model be trained small or distilled from a larger model?” | Training and compression recipe |
| Product engineer | “Is the compact model accurate enough for this task?” | Task-specific evaluation results |
- DistilBERT reduced model size and improved inference speed in its reported comparison.
- Sub-billion-parameter models can target mobile use cases.
- DistilBERT uses knowledge distillation during pretraining.
- Phi-1 uses selected textbook-quality data.
- Small language models can struggle to produce coherent and fluent text.
- TinyStories reports that its smaller model often repeats itself or makes no sense.
- TinyStories reports that one-layer models struggle substantially with following instructions.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMobileLLM Optimizing Sub-billion Parameter Language Models for On-Device Use Cases, Liu and colleagues · read 28 Sept 2026
- paperTinyStories: How Small Can Language Models Be and Still Speak Coherent English?, Eldan and Li · read 28 Sept 2026
- paperTextbooks Are All You Need, Gunasekar and colleagues · read 28 Sept 2026
- paperDistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, Sanh and colleagues · read 28 Sept 2026
- paperGemma: Open Models Based on Gemini Research and Technology, Gemma Team · read 28 Sept 2026