Dokumentationerreichbar
scikit-learn, Feature extraction
scikit-learn.org (externe Seite)
Die Dokumentation beschreibt im Wortlaut, wie aus einer Sammlung von Texten Zahlenreihen werden: zerlegen, zählen, gewichten. Sie benennt dabei die Eigenschaft, an der dieses Verfahren seine Grenze hat, nämlich dass die Stellung der Wörter im Text vollständig entfällt.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 05.09.2026:
We call vectorization the general process of turning a collection of text documents into numerical feature vectors.
bestätigt 24.09.2026Documents are described by word occurrences while completely ignoring the relative position information of the words in the document.
bestätigt 24.09.2026If we were to feed the direct count data directly to a classifier those very frequent terms would shadow the frequencies of rarer yet more interesting terms.
bestätigt 24.09.2026