NachschlagenQuellenregister

Dokumentationerreichbar

Qwen2.5-VL-7B-Instruct

huggingface.co (externe Seite)

Modellkarte des allgemeinen Bildmodells, das hier Handschrift und Diagramme übernimmt. Nach Herstellerangabe wertet es Texte, Diagramme, Symbole, Grafiken und Layouts in Bildern aus und gibt für Scans von Rechnungen, Formularen und Tabellen strukturierte Ergebnisse aus. Es kann Bildbereiche über Rahmen oder Punkte verorten. Die Familie umfasst Fassungen mit 3, 7 und 72 Milliarden Parametern.

geprüft 24.09.2026

Worauf sich diese Seite beruft, wörtlich, abgerufen am 07.09.2026:

  • it is highly capable of analyzing texts, charts, icons, graphics, and layouts within imagesbestätigt 24.09.2026
  • Generating structured outputs: for data like scans of invoices, forms, tables, etc. Qwen2.5-VL supports structured outputs of their contents, benefiting usages in finance, commerce, etc.bestätigt 24.09.2026
  • We have three models with 3, 7 and 72 billion parameters. This repo contains the instruction-tuned 7B Qwen2.5-VL model.bestätigt 24.09.2026

Wie eine Maschine ein Dokument zerlegtWo in der eigenen Texterkennung das Modell sitzt

Alle Quellen

Tippen Sie los.

↑↓ auswählenEnter öffnenDie Suche läuft im Browser. Nichts wird übertragen.