SchlüsselarbeitOriginalarbeiterreichbar
Architecture as Capability Equalizer for Coding Agents
Ein kontrollierter Versuch vom 22.08.2026 zur Frage, ob die Form einer Vorgabe das Ergebnis verändert. Fünf informationell gleichwertige Formate (formloser Text, Mermaid mit Bedingungen, OpenAPI, C4/Structurizr, TypeScript-Verträge) laufen über sechs Modelle dreier Hersteller, 90 Durchgänge. Der Befund kehrt die Erwartung um: Beim stärksten Modell macht die Form kaum einen Unterschied, beim schwächeren entscheidet sie, und code-nahe Formate holen den Abstand weitgehend auf.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 03.09.2026:
five informationally equivalent specification formats
bestätigt 24.09.2026Across 90 multi-turn agent trials, specification format shows a strong format x model interaction.
bestätigt 24.09.2026On the strongest models (Sonnet 4.6, GPT-5), format barely matters (quality spread 0.17-0.92)
bestätigt 24.09.2026On weaker models, format produces spreads of 0.83-2.42 points, with code-proximate formats (OpenAPI, TypeScript contracts) recovering most of the capability gap
bestätigt 24.09.2026TypeScript contracts triple API route coverage for the weakest model (33% to 100%).
bestätigt 24.09.2026Mid-tier models can consume more tokens than frontier models for worse output when they enter compilation debugging loops that stronger models avoid.
bestätigt 24.09.2026Self-validation rates collapse from 100% (Sonnet) to 0% (Gemini Flash) across the capability spectrum.
bestätigt 24.09.2026Structured architecture specifications serve as a capability equalizer, with value inversely proportional to model strength and the largest returns for cost-optimized deployments.
bestätigt 24.09.2026