Concepts

Lost in the middle

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

Language models often use facts near the opening or closing of a long input better than facts buried in the middle.

1 · What it is

“Lost in the middle” is the name for a habit of large language models, the AI systems behind chatbots. Give one a long input, and it tends to use facts near the start or the end better than facts in the middle. The name comes from a research paper by Nelson Liu and colleagues. It was first posted online in July 2023 and was published in the journal TACL.

A token is a small chunk of text, often part of a word. The context window is the most text, counted in tokens, that a model can read at once. A long-context model is one built to accept a very large window.

What the researchers did. They ran two tests. In the first, the model got a question plus a stack of short Wikipedia passages, each at most 100 tokens. Exactly one passage held the answer. The others were distractors, passages that look related but do not answer the question. Then the team moved the useful passage to different positions and measured accuracy each time. In the second test, the model got a long list of pairs, each a key (like a label) and a value. The list was written in JSON, a common data format. The model had to return the value for one key.

They tested open models, MPT-30B-Instruct and LongChat-13B (16K), and closed models, GPT-3.5-Turbo and Claude-1.3.

Two yardsticks. To have something to compare against, the team also ran two simple reference runs. In the “closed-book” run, the model got the question with no documents at all. It had to answer from what it learned in training, like a test with the book shut. In the “oracle” run, the model got only the one passage with the answer, and nothing else.

More documents, same pattern. The team tried stacks of 10, 20 and 30 documents. With 20 or 30 documents, the worst position scored lower than giving the model no documents at all.

A walk-through of the key-value test. Picture a long list of entries. Each key and each value is a UUID, a random string of letters and numbers such as a scrambled serial code. The team tried lists of 75, 140 and 300 pairs.

What they found. They drew a chart of accuracy at each position, and the line made a U shape. The authors call the edge at the start primacy bias (“primacy” means coming first). They call the edge at the end recency bias (“recency” means coming last). The dip appeared even for models built for long inputs.

One result stands out. With the answer placed in the middle, GPT-3.5-Turbo did worse on the question task than when it got no documents at all. Its no-document score was 56.1 percent. Burying the right passage hurt more than having nothing. On the key-value test, some models were perfect. Others struggled to find items in the middle.

A bigger window was not an easy fix. Versions of a model with a longer window often scored exactly the same as their shorter versions.

Why it matters. Many AI apps paste lots of text into one prompt, the message sent to the model. Some apps first run a search, then feed the documents it finds to the model. If the most useful document lands in the middle, the model may overlook it. The paper suggests reranking. That means sorting the found documents so the most useful ones go first.

The practical tip. Put the most important information near the start or the end. Keep the question near an edge too, not buried inside. Anthropic suggests placing long documents near the top and the question below them. It reports that putting the question last lifted answer quality as much as 30 percent in its tests. Google says a question usually works best at the end of a long prompt. OpenAI suggests repeating instructions both before and after the long context.

These tips reduce the risk but do not remove it. The size of the effect depends on the model and the task. The simplest check is the one the paper used: move the key fact around and see whether answers change.

2 · Why it exists

A model that accepts a long input does not always use all of it well.

Position changes answersMoving the one useful passage to a different spot can change how often the model answers correctly.
Middle is weakestAccuracy tends to be highest at the start or end and drops in the middle.
Bigger windows no cureModels with longer context windows were not necessarily better at using what they were given.
3 · How it works

Follow one test where only the position of the answer changes.

The researchers moved only the answer document and measured accuracy at each position.
  1. 1 · buildGive the model a question and several documents, where exactly one document holds the answer.
  2. 2 · moveRun the same test again with the answer document placed at different positions.
  3. 3 · measureRecord accuracy for each position and plot it as a curve.
  4. 4 · compareThe curve forms a U, highest at the start or end and lowest in the middle.

Practical rule: put the key information and the question at the start or the end, not buried in the middle.

4 · Where it's used
WhoWhat they askWhat it works with
Search-app builder“Which retrieved passages should go first in the prompt?”The order of retrieved documents
Prompt writer“Should my question go before or after a long document?”Where the query and instructions sit
Model evaluator“Does this model use evidence from every part of its input?”Accuracy at each position of the relevant text
5 · What it solves, and what it doesn't
solves
  • It names a measured weakness in how models use long inputs.
  • It gives a simple test for any model, by moving the relevant text and checking accuracy.
  • It explains why reranking retrieved documents, so the best ones come first, can help.
doesn't solve
  • It is not a fix; the middle of the input can still be weak.
  • A larger context window does not remove the effect on its own.
  • The exact size of the effect depends on the model and task.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperLost in the Middle: How Language Models Use Long Contexts, arXiv · read 28 Sept 2026
  2. paperLost in the Middle: How Language Models Use Long Contexts (full text PDF), arXiv · read 28 Sept 2026
  3. paperLost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics (ACL Anthology) · read 28 Sept 2026
  4. reponelson-liu/lost-in-the-middle, GitHub · read 28 Sept 2026
  5. docsPrompting best practices, Anthropic · read 28 Sept 2026
  6. docsLong context, Google AI for Developers · read 28 Sept 2026
  7. docsGPT-4.1 prompting guide, OpenAI · read 28 Sept 2026