Using LLMs Well: Context, Temperature, Hallucinations And Retrieval
Why language models invent facts, how settings like temperature change their output, and how retrieval-augmented generation grounds answers in your own documents.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
Many of the problems people hit with language models in real work trace back to a few common causes: the model never saw the information it needed, it was given too much to read at once, its settings encouraged variety where consistency was needed, or it guessed instead of saying it did not know.
This article explains four practical levers that address those causes: the context window, sampling settings such as temperature, the research on why models hallucinate, and retrieval-augmented generation (RAG). It assumes you know the next-token loop described in How An LLM Reads And Writes.
The Context Window: What The Model Can See
A model has two broad sources of information when it answers. One is what it absorbed into its parameters during training, which mostly stops at a cut-off date. The other is the context window: the instructions, conversation history, documents, search or tool results and question supplied with the request, all measured in tokens. Anything outside those two sources is invisible to the model. Researchers describe these as parametric memory, stored in the model’s weights, and non-parametric memory, supplied from outside.1 The cut-off is not an absolute wall: later fine-tuning can teach a model new facts, although one 2024 study found it learned them slowly and became more prone to hallucinate as it did.2
Context windows have grown a great deal, and limits vary widely between products and change often, so check your provider’s documentation rather than relying on a remembered figure. More room does not guarantee that the model reads all of it with equal care. In a 2023 study, Liu and colleagues placed the one document that answered a question at different positions among many unhelpful ones. The models they tested tended to answer best when that document came first or last, and their accuracy dropped when it sat in the middle. Models built for long inputs showed the same dip.3 Newer models may behave differently, so test the one you use. The safe habits are to send the model what it needs, not everything you have, and to put the most important material near the beginning or end.
Temperature And Sampling: How Much Variety
At each step the model produces a probability for every possible next token. Decoding settings decide how a token is chosen from that distribution.
Temperature reshapes the distribution before a choice is made: for any temperature above zero, the model’s raw scores are divided by the temperature before being turned into probabilities. A setting of zero is usually treated as a special case that always picks the highest-scoring token. A low temperature sharpens the distribution, so the most likely tokens become even more likely and output becomes more predictable. A high temperature flattens it, giving less likely tokens a better chance and making output more varied. Lowering the temperature tends to improve quality at the cost of variety.4 Top-p, or nucleus, sampling, introduced by Holtzman and colleagues, takes a different route: it samples only from the smallest set of tokens whose combined probability reaches a threshold, cutting off the long tail of unlikely options.4 Their work showed that the decoding method alone has a large effect on the quality of text from the same model.
| Task | Suggested Direction | Why |
|---|---|---|
| Extracting fields from an invoice | Low temperature | You want answers to vary as little as possible |
| Classifying support tickets | Low temperature | More consistent results are easier to audit |
| Drafting marketing headlines | Higher temperature | Variety gives more options to choose from |
| Summarising a policy | Low to moderate | Faithfulness matters more than flair |
These controls are not always available. As of October 2026, Anthropic’s API documentation says its models released after Claude Opus 4.6 no longer accept a temperature setting other than the default, so check what your provider allows.5
Why Models Hallucinate
A hallucination is output that sounds plausible but is false or unsupported, such as a made-up citation, a non-existent policy clause or a wrong figure. A 2025 paper by Kalai, Nachum, Vempala and Zhang, researchers at OpenAI and Georgia Tech, offers one research-based explanation.6
They argue the problem has two roots. First, during pretraining some errors are statistically unavoidable when true and false statements are hard to tell apart from the data, for example facts such as a person’s birthday that appear only once in the training data. Second, and more fixable, most benchmarks score answers simply as right or wrong, so a model that guesses when unsure scores better than one that admits uncertainty. The authors put it bluntly: models are “optimized to be good test-takers”.6 They propose changing how existing mainstream benchmarks are scored, for example by telling the model that wrong answers cost points while abstaining costs nothing, so that honest uncertainty is no longer penalised.
For an organisation, the useful question follows directly: does your system allow, and reward, the model saying it does not know? Measure correct answers, wrong answers and abstentions separately rather than as one accuracy number.
Retrieval-Augmented Generation: Bringing In The Facts
RAG addresses the “never saw the information” problem. Instead of relying on what the model memorised, the system searches a trusted collection of documents for passages relevant to the question and supplies them in the prompt. The idea was set out by Lewis and colleagues in 2020, who combined a language model’s built-in, parametric memory with a searchable index of Wikipedia passages. They reported more specific and factual output than a model without retrieval, and noted two further benefits: answers can point to their sources, and knowledge can be updated by changing the documents rather than retraining the model.1
- Prepare The Documents
Split policies, contracts or manuals into passages and store an embedding of each in a vector index.
- Embed The Question
Convert the user question into an embedding using the same embedding model.
- Search The Index
Find the passages whose embeddings are closest in meaning to the question.
- Select The Top Passages
Keep a small number of the most relevant passages, ideally with their source details.
- Assemble The Prompt
Combine instructions, the passages and the question, and tell the model to answer only from the passages.
- Generate A Cited Answer
The model answers and points to the passages that support each claim.
- Evaluate
Check both whether the right passages were found and whether the answer stays faithful to them.
RAG improves grounding but does not guarantee it. Search can return the wrong passage, the right passage can be lost among too many others, and the model can still add claims the passages do not support. Retrieved documents are also a route for attack: text planted in a document can carry instructions the model may follow. Greshake and colleagues demonstrated this in 2023 and named it indirect prompt injection;7 it is covered further in AI Security. When RAG feeds a system that takes actions, see Agentic AI Security.
Evaluating A RAG System
Because RAG has two halves, it needs to be tested in two halves. The RAGAS framework, published in 2023, proposes automated checks that do not require a human-written reference answer for every question.8 It asks three questions: did the search bring back passages that bear on the question without much padding, can every claim in the answer be traced to those passages, and does the answer respond to what was actually asked?
In practice, keep a set of real questions with known good sources. Check retrieval on its own: did the right passage appear in the results? Then check generation: is every claim in the answer supported by a retrieved passage? Track abstentions as well, since an answer of “the documents do not say” is often the correct one.
Footnotes
-
P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, NeurIPS 2020, arXiv:2005.11401. arxiv.org ↩ ↩2
-
Z. Gekhman et al., “Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?”, EMNLP 2024, arXiv:2405.05904. arxiv.org ↩
-
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni and P. Liang, “Lost in the Middle: How Language Models Use Long Contexts”, Transactions of the Association for Computational Linguistics, arXiv:2307.03172, July 2023. arxiv.org ↩
-
A. Holtzman, J. Buys, L. Du, M. Forbes and Y. Choi, “The Curious Case of Neural Text Degeneration”, ICLR 2020, arXiv:1904.09751. arxiv.org ↩ ↩2
-
Anthropic (vendor documentation), “Create a Message” API reference, temperature parameter, accessed 7 October 2026. platform.claude.com ↩
-
A. T. Kalai, O. Nachum, S. S. Vempala and E. Zhang, “Why Language Models Hallucinate”, arXiv:2509.04664, 4 September 2025. arxiv.org ↩ ↩2
-
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz and M. Fritz, “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection”, arXiv:2302.12173, February 2023. arxiv.org ↩
-
S. Es, J. James, L. Espinosa-Anke and S. Schockaert, “Ragas: Automated Evaluation of Retrieval Augmented Generation”, arXiv:2309.15217, September 2023. arxiv.org ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.