Konzeption und Implementierung eines Systems für Visual Document Retrieval mittels Vision-Language-Modellen zur wissensbasierten Informationsgewinnung

K. G., 2026

A system for visual document retrieval based on vision‑language models was designed and evaluated against traditional OCR‑driven pipelines. In a reproducible benchmark that compared five distinct retrieval paths—including text‑only, hybrid, and pure visual approaches—the native visual methods Ops‑ColQwen3 and Gemini Embedding 2 achieved full top‑5 coverage and superior ranking on complex PDF queries, while also revealing higher memory and model‑dependency costs. The findings support the view that visual‑only retrieval can markedly improve search in layout‑heavy documents, provided that resource constraints are taken into account.