Dokumentationerreichbar
Vision
platform.claude.com (externe Seite)
Die Herstellerdokumentation zum Bild als Eingabe, und die einzige der drei, die ihre Abrechnung als Formel ausschreibt: Ein Bild wird in Blöcke von 28 mal 28 Pixeln zerlegt, und jeder Block ist ein sichtbares Token. Daraus folgt die Zahl über die Fläche, aufgerundet in beide Richtungen, gedeckelt bei 1568 sichtbaren Token für die ältere Stufe und 4784 für die neuere. Der Abschnitt Limitations zählt sieben Grenzen auf, darunter die für dieses Projekt wichtigste: Ob ein Bild von einer Maschine stammt, kann das Modell nicht sagen.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 07.08.2026:
Claude views images in patches instead of pixels. Each patch is a 28×28-pixel block of the image, referred to as a visual token.
bestätigt 24.09.2026Images larger than either limit are downscaled before processing
bestätigt 24.09.2026When an image is downsized, Claude scales it to the largest size that fits the tier's limits while preserving its aspect ratio.
bestätigt 24.09.2026High-resolution images can use up to roughly three times more visual tokens than the same image on a standard-tier model.
bestätigt 24.09.2026Animations are unsupported, and only the first frame is used.
bestätigt 24.09.2026Claude works best when images come before text.
bestätigt 24.09.2026Claude cannot determine whether an image is AI-generated and might be incorrect if asked.
bestätigt 24.09.2026Claude can give approximate counts of objects in an image but might not always be precisely accurate, especially with large numbers of small objects.
bestätigt 24.09.2026Claude's coordinate and localization outputs are approximate.
bestätigt 24.09.2026Although Claude can analyze general medical images, it is not designed to interpret complex diagnostic scans such as CTs or MRIs.
bestätigt 24.09.2026Claude might hallucinate or make mistakes when interpreting low-quality, rotated, or very small images under 200 pixels.
bestätigt 24.09.2026
Multimodal im GlossarVisuelles Token im Glossar13 Was ein Modell auf einem Bild erkennt