Does retrieval augmentation reduce hallucination?
rag hallucination lit
Structured projection of the notebook.
Skeleton — no draftThe notebook has no draft (
drafts/current or drafts/vN), so this page is an honest structured projection of the notebook: its open questions, its claims grouped by verification status, and its sources by grade. Write the narrative as a flip draft and regenerate to replace this.Questions
- [Q1] Do retrieval-augmentation factuality gains persist for post-2024 frontier models, whose parametric knowledge is far stronger than the 2022-2023 models these st… — open
Claims by status
verified
- [C15] Across all three included studies, retrieval interventions improved every reported hallucination-adjacent metric against non-retrieval baselines - but each stu… (load-bearing)
asserted
- [C11] A 2.7B-parameter LM with retrieval (GPT-Neo + Contriever) outperformed vanilla GPT-3 175B on PopQA long-tail factual QA (load-bearing)
- [C12] Retrieval augmentation, not scale, was the main contributor to Atlas's superior factual consistency over closed-book LLaMA; scaling alone yielded roughly 1-3 p… (load-bearing)
- [C13] Post-hoc retrieval-augmented correction improved long-form factuality metrics by up to 30 percent over prior correction methods
- [C14] Retrieval can hurt: on roughly 10 percent of PopQA questions, augmentation degraded accuracy because the retrieved text itself was poor
Sources by grade
Grade A
- [A1] arxiv.org/pdf/2311.01307 (independent)
- [A2] arxiv.org/pdf/2212.10511 (independent)
- [A3] arxiv.org/pdf/2410.15667 (independent)