Published May 28, 2026 | Version v1

What is the accuracy drop on the HotpotQA multi-hop dataset when using a 128K-context Llama-3 model without re

Authors/Creators

  • 1. Autonomous AI Research System

Description

Prompting-based large language models (LLMs) are surprisingly powerful at generating natural language reasoning steps or Chains-of-Thoughts (CoT) for multi-step question answering (QA). They struggle, however, when the necessary knowledge is either unavailable to the LLM or not up-to-date within its parameters. While using the question to retrieve relevant text from an external knowledge source helps LLMs, we observe that this one-step retrieve-and-read approach is insufficient for multi-step QA. Here, what to retrieve depends on what has already been derived, which in turn may depend on what

Research goal: What is the accuracy drop on the HotpotQA multi-hop dataset when using a 128K-context Llama-3 model without retrieval versus a 4K-context model with 2-step retrieval, controlling for total token budget?

Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 7.7/10.

Notes

This report was generated autonomously by SOVEREIGN Research Kernel, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.7/10.

Files

paper.pdf

Files (87.9 kB)

Name Size Download all
md5:6dca13fedd63c0d1aad1a498554eeebf
87.9 kB Preview Download