Why LLMs struggle to find the information in the middle of long documents?

A Billion-Dollar Problem which if not solved makes it difficult for LLMs to scale beyond where they are today

Why LLMs struggle to find the information in the middle of long documents?

Imagine paying premium prices for a cutting-edge AI system that claims to process 200,000 tokens of text—roughly equivalent to a 400-page document—only to discover it systematically fails to find crucial information buried in the middle pages. This isn't a hypothetical scenario; it's the reality facing organisations deploying today's most advanced large language models (LLMs).

The phenomenon, known as "lost in the middle," was first systematically documented by Liu et al. (2023) and reveals a fundamental flaw that undermines the practical utility of extended-context AI systems. Despite impressive claims about processing hundreds of thousands of tokens, these models exhibit severe position bias—performing optimally when relevant information appears at the beginning or end of their input while showing dramatic performance degradation for information positioned in middle segments.

After diving deep into this emerging research area, I've spent months analyzing the growing body of literature to understand the full scope and implications of this problem. The result is a comprehensive research synthesis that reveals just how pervasive and costly this limitation has become.

The Scope of the Problem: Universal and Costly

Through my extensive analysis of research spanning 2023-2025, I've found that this isn't an isolated bug affecting specific models—it's a universal phenomenon rooted in the fundamental architecture of transformer-based language models. The implications extend far beyond academic curiosity into real-world applications where organizations depend on AI for critical decision-making.

Consider these scenarios:

  • A legal AI analyzing a 100-page contract where crucial terms appear on page 50
  • A medical AI reviewing patient histories where diagnostic information is scattered throughout extensive records
  • An enterprise knowledge system processing regulatory documents where compliance requirements are distributed across lengthy texts

In each case, current long-context models may systematically fail to identify critical information based purely on its positional placement rather than its semantic importance.

The Economic Reality: Paying More for Less

The economic implications are staggering. Extending context from 4,000 to 128,000 tokens typically increases inference costs by an order of magnitude due to quadratic scaling in attention mechanisms. Yet research reveals that models effectively utilize only 10-20% of their claimed context capacity. This represents what the paper calls "a fundamental market inefficiency where customers pay for capabilities that models cannot actually deliver."

The computational waste is enormous. Organisations investing in long-context capabilities expect proportional improvements in processing effectiveness, yet position bias ensures diminishing returns as context length increases. Companies like Google claiming 10 million token capabilities and Anthropic achieving 200,000 tokens may be offering capabilities that provide minimal practical value beyond much shorter contexts.

The Six Pillars of a Fundamental Limitation

1. Architectural Origins: Built-In Bias

The position bias phenomenon emerges from fundamental design choices in transformer architecture. The causal masking required for autoregressive generation creates an asymmetric information flow that inherently favors beginning positions. When combined with positional encodings, this creates what researchers call "systematic disadvantage for middle positions that cannot be easily remedied through simple architectural modifications."

Recent graph-theoretic analysis by Wu et al. (2025) provides mathematical proof that this bias is baked into the architecture itself, not just an implementation quirk that can be easily fixed.

2. Universal Empirical Evidence

The seminal experiments by Liu et al. demonstrated that GPT-3.5-Turbo's performance on multi-document questions actually dropped below its zero-shot baseline when relevant information was positioned in middle sections—meaning additional context harmed rather than helped performance.

This finding has been confirmed across multiple evaluation frameworks:

The universality extends beyond text processing. Multimodal research shows that visual-language models exhibit similar positional biases when processing image sequences, suggesting the problem transcends individual modalities.

3. Computational Inefficiency at Scale

The paper reveals massive computational waste in current deployments. While organizations pay premium prices for extended context processing, the majority of this computation provides minimal value. The comprehensive survey by Dong et al. highlights how this inefficiency affects the sustainability and scalability of AI deployments.

4. Limited Mitigation Strategies

Current solutions fall short of comprehensive resolution:

  • Query repositioning: Placing queries at beginning and end positions helps but doesn't solve complex reasoning tasks
  • Attention strengthening: The "never lost in the middle" approach by He et al. shows modest improvements but requires substantial architectural modifications
  • Advanced prompting: Multiple inference passes can help but effectively double computational costs

5. Theoretical Understanding

Graph-theoretic frameworks provide mathematical foundations explaining why position bias intensifies with increased context length. Tokens in middle positions suffer from reduced connectivity in the attention graph, receiving less direct attention and having fewer pathways for information aggregation compared to edge tokens.

6. Promising Future Directions

The most promising solutions involve fundamental architectural reconceptualisation:

  • Memory-augmented processing: Infini-attention by Munkhdalai et al. demonstrates effectiveness on 1M sequence length tasks
  • Multi-agent collaboration: Chain of Agents frameworks distribute reasoning across specialised components
  • Alternative architectures: State-space models like Mamba offer linear computational complexity and potentially more balanced information processing

Real-World Impact: Beyond Academic Interest

The position bias problem affects organisations across industries:

Legal Technology: Law firms using AI for contract analysis may miss critical clauses based purely on their document position. This creates liability risks and undermines the reliability of AI-assisted legal review.

Healthcare Systems: Medical AI processing patient records may overlook diagnostic information depending on where it appears in lengthy medical histories, potentially affecting patient care quality.

Financial Services: Regulatory compliance systems may fail to identify important requirements buried in middle sections of complex regulatory documents, creating compliance risks.

Enterprise Knowledge Management: Companies deploying AI for internal document processing may experience systematic blind spots based purely on information positioning rather than relevance.

The Research Opportunity

For the academic community, this phenomenon opens rich research avenues spanning theoretical analysis, architectural innovation, and evaluation methodology development. The mathematical frameworks emerging from attention analysis provide principled foundations for designing solutions, while the documented universality across model families suggests breakthrough solutions could have broad applicability.

Key research directions include:

  • Novel attention mechanisms that balance global and local information processing
  • Training methodologies that explicitly counteract positional asymmetries
  • Evaluation frameworks that distinguish genuine capabilities from positional artefacts
  • Deployment strategies that optimise information presentation for current model limitations

The Commercial Opportunity

For AI practitioners and technology leaders, the position bias problem represents both a significant challenge and a substantial market opportunity. Organizations that develop effective solutions could achieve competitive advantages in applications requiring reliable long-context processing.

The economic analysis reveals substantial computational inefficiencies that create clear opportunities for more effective solutions. Early recognition of these limitations enables more informed deployment decisions and risk management strategies.

Current Industry Response

The research community has responded with increasingly sophisticated evaluation frameworks and mitigation strategies. The LongBench evaluation suite provides bilingual, multitask assessment capabilities, while efficient training approaches like LongRecipe demonstrate that context extension can be achieved with reduced computational overhead—though fundamental utilisation limitations persist.

Recent work on self-improvement approaches suggests models can potentially overcome limitations through iterative refinement, though computational overhead remains a concern for practical deployment.

What This Means for Your Organisation

If your organisation is considering or currently deploying long-context AI systems, this research has immediate practical implications:

  1. Audit current deployments: Evaluate whether critical information might be systematically missed based on document positioning
  2. Implement position-aware strategies: Consider document restructuring to place important information at beginning or end positions
  3. Budget realistically: Understanding that only a fraction of claimed context capacity is effectively utilised affects cost-benefit calculations
  4. Monitor for blind spots: Develop testing procedures that specifically evaluate middle-position performance in your use cases

The Path Forward

The research suggests that addressing position bias will likely require moving beyond traditional single-model architectures toward more sophisticated multi-component systems. The most promising approaches involve:

  • Retrieval-augmented systems that dynamically restructure context based on relevance rather than position
  • Hybrid architectures that combine traditional transformers with specialised components designed to overcome specific limitations
  • Memory-augmented processing that maintains efficient access to historical information without traditional attention limitations

Conclusion: A Critical Juncture for AI Development

The lost in the middle phenomenon represents more than a technical limitation—it's a fundamental constraint that affects the reliability, efficiency, and economic viability of current AI deployments. Understanding this problem is crucial for anyone working with or investing in long-context AI systems.

My analysis of the emerging research reveals both the depth of the challenge and the promise of potential solutions. Organisations that recognise and address these limitations early will be better positioned to deploy AI effectively while avoiding the systematic blind spots that affect current implementations.

For researchers, practitioners, and business leaders alike, this represents new opportunities for innovation and competitive advantage. The next breakthrough in AI may come not from scaling existing architectures but from fundamental innovations that address the systematic limitations revealed by position bias research.


Want the complete technical analysis? After extensive research across dozens of recent papers, I've synthesised the current state of knowledge into a comprehensive 25 -page research paper that provides deep analysis that you may find useful for your work.

Download the paper below for the full technical details. It's free.

My paper includes:

  • Detailed mathematical analysis of why position bias occurs at the architectural level
  • Complete evaluation methodology review across all major benchmarking frameworks
  • Comprehensive mitigation strategy analysis with practical implementation guidance
  • Six-pillar framework for understanding every aspect of the position bias problem
  • Complete bibliography with direct links to all 40+ papers analysed
  • Technical implementation strategies for organisations looking to address these limitations
  • Future research directions and commercial opportunities

This synthesis represents months of analysis across the cutting-edge research in long-context AI systems—essential reading for anyone serious about understanding and solving one of AI's most critical current limitations.

Key papers referenced in this analysis:

For organisations interested in implementing solutions or researchers exploring this phenomenon, the full paper provides the technical depth and comprehensive analysis needed to understand and address this critical limitation in current AI systems.