Concepts

Small language models

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Small-model research includes sub-billion-parameter language models for mobile deployment.

1 · What it is

MobileLLM studies sub-billion-parameter models for mobile deployment. TinyStories trains models below ten million parameters.

Phi-1 has 1.3 billion parameters and uses selected textbook-quality data. DistilBERT uses knowledge distillation and reduced BERT model size by 40 percent while retaining 97 percent of its reported language-understanding capability.

MobileLLM emphasizes model architecture at the sub-billion scale. The Gemma paper includes 2 billion and 7 billion parameter models.

2 · Why it exists

Large pretrained models can be difficult to run under constrained compute budgets.

Device limitsMobileLLM focuses on language models with fewer than one billion parameters for mobile deployment.
Serving costDistilBERT was designed as a smaller, faster and lighter model.
Data mattersPhi-1 has 1.3 billion parameters.
3 · How it works

Compare parameter scale with the deployment budget.

MobileLLM explicitly studies sub-billion-parameter models for on-device use cases.
  1. 1 · boundSet the deployment resource envelope.
  2. 2 · sizeChoose a parameter count that fits that envelope.
  3. 3 · compareModel architecture matters at sub-billion scale.
  4. 4 · measureRun a comparative on-device study.

TinyStories studies models below ten million parameters; phi-1 has 1.3 billion parameters.

4 · Where it's used
WhoWhat they askWhat it works with
Mobile developer“Can this model run within the device budget?”Parameters, memory and measured device latency
Model trainer“Should the model be trained small or distilled from a larger model?”Training and compression recipe
Product engineer“Is the compact model accurate enough for this task?”Task-specific evaluation results
5 · What it solves, and what it doesn't
solves
  • DistilBERT reduced model size and improved inference speed in its reported comparison.
  • Sub-billion-parameter models can target mobile use cases.
  • DistilBERT uses knowledge distillation during pretraining.
  • Phi-1 uses selected textbook-quality data.
doesn't solve
  • Small language models can struggle to produce coherent and fluent text.
  • TinyStories reports that its smaller model often repeats itself or makes no sense.
  • TinyStories reports that one-layer models struggle substantially with following instructions.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperMobileLLM Optimizing Sub-billion Parameter Language Models for On-Device Use Cases, Liu and colleagues · read 28 Sept 2026
  2. paperTinyStories: How Small Can Language Models Be and Still Speak Coherent English?, Eldan and Li · read 28 Sept 2026
  3. paperTextbooks Are All You Need, Gunasekar and colleagues · read 28 Sept 2026
  4. paperDistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, Sanh and colleagues · read 28 Sept 2026
  5. paperGemma: Open Models Based on Gemini Research and Technology, Gemma Team · read 28 Sept 2026