What is the accuracy drop on the HotpotQA multi-hop dataset when using a 128K-context Llama-3 model without re
Description
Prompting-based large language models (LLMs) are surprisingly powerful at generating natural language reasoning steps or Chains-of-Thoughts (CoT) for multi-step question answering (QA). They struggle, however, when the necessary knowledge is either unavailable to the LLM or not up-to-date within its parameters. While using the question to retrieve relevant text from an external knowledge source helps LLMs, we observe that this one-step retrieve-and-read approach is insufficient for multi-step QA. Here, what to retrieve depends on what has already been derived, which in turn may depend on what
Research goal: What is the accuracy drop on the HotpotQA multi-hop dataset when using a 128K-context Llama-3 model without retrieval versus a 4K-context model with 2-step retrieval, controlling for total token budget?
Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 7.7/10.
Notes
Files
paper.pdf
Files
(87.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:6dca13fedd63c0d1aad1a498554eeebf
|
87.9 kB | Preview Download |