• We study how pretraining and retrieval interact across OLMo-2-based language models with 30M to 3B parameters, varying training exposure, datastore size, and whether retrieved data was previously seen. Retrieval benefits depend on the evaluation metric: smaller models improve more in gold-answer perplexity, while larger, more-pretrained models gain more in accuracy. The results support treating external retrieval and parametric knowledge as complementary resources whose value depends on the task and training regime.

    Karan Singh, Michael Yu, Varun Gangal, Zhuofu Tao, Sachin Kumar, Emmy Liu, Steven Y. Feng
    arXiv · August 7, 2026 · projects: scaling-laws
  • Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs. Why is this? We review existing definitions of hallucination and fold them into a single, unified definition wherein prior definitions are subsumed. We argue that hallucination can be unified by defining it as simply inaccurate (internal) world modeling, in a form where it is observable to the user. For example, stating a fact which contradicts a knowledge base OR producing a summary which contradicts the source. By varying the reference world model and conflict policy, our framework unifies prior definitions. We argue that this unified view is useful because it forces evaluations to clarify their assumed reference 'world', distinguishes true hallucinations from planning or reward errors, and provides a common language for comparison across benchmarks and discussion of mitigation strategies. Building on this definition, we connect the framework to HalluWorld, a complementary benchmark using fully specified reference world models to study model hallucinations.

    Emmy Liu, Varun Gangal, Chelsea Zou, Michael Yu, Xiaoqi Huang, Alex Chang, Zhuofu Tao, Karan Singh, Sachin Kumar, Steven Y. Feng
    ICML 2026 · June 12, 2026 · projects: hallu-world
  • HalluWorld measures hallucination against explicit reference worlds in gridworlds, chess, and realistic terminal tasks. Controlled observations and automatically generated labels make it possible to isolate errors in perception, state tracking, and causal simulation. Evaluations of frontier and open-weight models show that tracking changes and predicting consequences remain difficult, even with extended thinking, while terminal tasks also expose weaknesses in abstention.

    Emmy Liu, Varun Gangal, Michael Yu, Zhuofu Tao, Karan Singh, Sachin Kumar, Steven Y. Feng
    arXiv · May 19, 2026 · projects: hallu-world