
Datum: 23.09.2026 bis 24.09.2026
Uhrzeit: 10:00 bis 16:30 Uhr
Ort: Leibniz-Institut für Europäische Geschichte (IEG) | Alte Universitätsstraße 19 | D - 55116 Mainz
Diese Veranstaltung wird organisiert durch das
Format:
Bring-your-own-data-Lab
Build Your Own HistoRAG
Source sovereignty and historical method when working with LLMs
Workshop led by:
Noah Kim-Baumann und Aurel Daugs (HU Berlin)
Description:
Large language models are changing how researchers work with large source corpora, but the standard Retrieval-Augmented Generation (RAG) systems behind them are built for general-purpose, consumer-facing applications, often centred on factual question-answering. They treat similarity as relevance, obscure how sources are selected, flatten change over time, and present their output as answers rather than as something to be questioned. Those assumptions sit awkwardly with interpretive, source-critical scholarship.
This lab takes those problems as its starting point. Working with HistoRAG, a prototype developed at the Chair of Digital History at the Humboldt-Universität zu Berlin that builds historical method into the RAG process itself, we treat retrieval-augmented reading as a research process the scholar controls rather than a black box. Participants bring their own corpus, sent in advance, and over two days we work on and test a minimal HistoRAG setup against it, opening up the decisions that shape every result and watching how each choice changes what the system returns. It is an exploratory lab, focussing on surfacing open questions, needs, and limits of humanist-driven LLM work. The goal is to develop an understanding of source-sovereign, source-critical work with LLMs that keeps interpretive control and epistemic agency with the researcher.
The workshop is aimed at researchers across the humanities, cultural, and social sciences who work with their own large text corpora. We take the perspective that the conveniences LLMs offer (natural-language access, seamless answers, agentic decision-making) quietly hide the decisions that constitute scholarship and outsource the researcher’s epistemic agency. Reclaiming that agency is what this lab is about. The two days combine short impulse talks, hands-on building, and one-to-one mentoring on your own data. No programming experience is required, and corpora should be machine-readable. The lab is held in English and limited to 15 participants.
Your data:
Participants bring their own machine-readable text corpus. Please send your corpus by 13 September 2026 so that it can be prepared in advance of the lab.
If you cannot bring your own material, you can work with one of three prepared corpora instead:
- UN General Debate Corpus (English)
- Der Spiegel (German)
- Bundestagsprotokolle (German)
Data processing and privacy:
During the lab, model interactions can run on an open-weights model hosted on our own servers at the Humboldt-Universität zu Berlin. Your corpus is not passed to any external or commercial provider, and it is deleted from our servers after the lab. The system can also be connected to external cloud-based models with an API key, but this is optional and not needed in order to take part.
Requirements:
- Researchers in the humanities, cultural and social sciences working with their own large text corpora
- No programming experience required, as all work during the lab is browser-based
- Corpora must be machine-readable, ideally with metadata
- Own laptop
Contact and Registration:
Johanna Mauermann
hermes@ieg-mainz.de
Please provide the following information when registering:
- Your area of expertise
- What experience do you have with Retrieval-Augmented Generation (RAG) systems?
- What kind of data are you bringing with you?
The lab is held in English and limited to 15 participants.
Registration deadline: 13 September 2026
Programme
Day 1: Wednesday, 23 September 2026
10:00 Registration and coffee
10:30 Welcome and introductions
10:45 Discussion 1: How we work with AI tools, and the usability paradox
11:15 Coffee break
11:30 Impulse 1: LLM fundamentals and the settings we control
12:30 Lunch
13:30 Hands-on 1: A corpus in a consumer RAG tool
14:15 Coffee break
14:30 Impulse 2: RAG fundamentals
15:30 Impulse 3: HistoRAG, a source-critical RAG architecture
16:00 Hands-on 2: First exploration of your own corpus
17:00 Wrap-up and outlook
18:00 Workshop dinner (at own expense)
Day 2: Thursday, 24 September 2026
9:30 Welcome back and recap
9:45 Discussion 2: What counts as relevant in your project, and can you write it down?
10:30 Coffee break
10:45 Impulse 4: Evaluating retrieval with LLM-as-judge
11:30 Hands-on 3: Running LLM-as-judge with your own relevance criteria
12:30 Lunch
13:30 Hands-on 4 and mentoring: Analysis, Zwischentexte and citation tracing
14:45 Coffee break
15:00 Discussion 3: The role of LLM-produced texts in research
15:45 Plenary Discussion on maintaining Epistemic Agency in LLM-assisted Research Workflows
16:15 Wrap-up, next steps and feedback
16:30 Close