Image-based Model Categorization tool - Categorising 3D heritage models with locally run vision-capable LLMs
Résumé fourni par la source
Webinar presenting the Image-based Model Categorization tool developed at PCSS within 3DBigDataSpace, which assigns semantic categories to 3D heritage models from their rendered appearance. Categorisation is powered by vision-capable large language models that accept both text and images; the tool uses open-source models (such as MiniCPM-V) run locally via Ollama, so it is free and independent of commercial LLM providers. Model responses are constrained to a limited set of terms mapped to categories, and each 3D model is prompted several times to counter hallucinations. Rendering is done in Blender (Eevee, flat lighting and ambient occlusion), normalising each model and producing 12 orbiting orthographic views. A single Go application ties rendering, classification and result collection into one pipeline, shipped as a Docker container with optional NVIDIA GPU acceleration. A demo dataset of example models is included as a second file in this record.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.