SchlüsselarbeitOriginalarbeiterreichbar
Stop Comparing LLM Agents Without Disclosing the Harness
Ein Positionspapier vom 07.05.2026, das die Umgebung zur eigentlichen Stellgröße erklärt. Es benennt die Bestandteile eines Harness einzeln (Kontextaufbau, Tool-Verkehr, Ablaufsteuerung, Prüfung) und hält fest, dass heutige Vergleiche Verbesserungen der Umgebung dem Modell zuschreiben. Die stärkste verfügbare Belegstelle für die Frage, wie viel der Aufbau um ein Modell herum ausmacht.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 29.08.2026:
the agent execution harness, namely the infrastructure layer that governs context construction, tool interaction, orchestration, and verification around a language model, is often a stronger determinant of agent performance than the model it wraps
bestätigt 24.09.2026performance variance is governed more by harness configuration than by model choice
bestätigt 24.09.2026small harness changes can produce performance shifts that exceed those obtained by substituting one model for another
bestätigt 24.09.2026current evaluation protocols therefore systematically misattribute harness-level gains to model improvements
bestätigt 24.09.2026Until harness specifications are disclosed, leaderboard comparisons for long-horizon agents should be treated as incomplete and potentially misleading
bestätigt 24.09.2026