Ankündigung des Herstellerserreichbar
Anthropic, Introducing Contextual Retrieval
Anthropic beschreibt am 19.09.2024 einen vollständigen Suchaufbau und beziffert jede Stufe einzeln. Der Kern des Verfahrens ist ein erklärender Vorspann, den ein Modell für jedes Textstück aus dem ganzen Dokument erzeugt; darauf setzen Wortsuche neben Ähnlichkeitssuche und ein zweiter Sortierdurchgang auf. Die Zahlen sind Herstellerangaben aus einem eigenen Versuchsaufbau, sie nennen jeweils den Ausgangswert mit und sind damit nachrechenbar. Wer sie ohne den Vorspann liest, schreibt die Wirkung der falschen Stufe zu. Der Aufsatz gibt zugleich die Größenordnungen, mit denen solche Aufbauten arbeiten, und nennt die Grenze, unterhalb derer sich der ganze Aufwand erübrigt.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 06.08.2026:
By leveraging both BM25 and embedding models, traditional RAG systems can provide more comprehensive and accurate results, balancing precise term matching with broader semantic understanding.
bestätigt 24.09.2026Combining Contextual Embeddings and Contextual BM25 reduced the top-20-chunk retrieval failure rate by 49%
bestätigt 24.09.2026Reranked Contextual Embedding and Contextual BM25 reduced the top-20-chunk retrieval failure rate by 67%
bestätigt 24.09.2026instructs the model to provide concise, chunk-specific context that explains the chunk using the context of the overall document
bestätigt 24.09.2026The resulting contextual text, usually 50-100 tokens, is prepended to the chunk before embedding it and before creating the BM25 index.
bestätigt 24.09.2026Contextual Embeddings reduced the top-20-chunk retrieval failure rate by 35%
bestätigt 24.09.2026We use 1 minus recall@20 as our evaluation metric, which measures the percentage of relevant documents that fail to be retrieved within the top 20 chunks.
bestätigt 24.09.2026Adding more chunks into the context window increases the chances that you include the relevant information. However, more information can be distracting for models so there's a limit to this.
bestätigt 24.09.2026If your knowledge base is smaller than 200,000 tokens (about 500 pages of material), you can just include the entire knowledge base in the prompt that you give the model, with no need for RAG or similar methods.
bestätigt 24.09.2026BM25 refines this by considering document length and applying a saturation function to term frequency
bestätigt 24.09.2026Combine and deduplicate results from (3) and (4) using rank fusion techniques
bestätigt 24.09.2026