Yugong09/GeoGuess
VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Apr 19, 2026Architecture:Transformer Featherless Exclusive Cold
Yugong09/GeoGuess is a 4 billion parameter vision-language model based on the Qwen3-VL-4B-Instruct architecture, developed by Yugong09. This model is designed for multimodal tasks, leveraging its vision capabilities to process and understand image inputs. With a context length of 32768 tokens, it excels in applications requiring both visual comprehension and natural language interaction.
Loading preview...
Yugong09/GeoGuess: A Vision-Language Model
Yugong09/GeoGuess is a 4 billion parameter vision-language model built upon the Qwen3-VL-4B-Instruct architecture. This model integrates robust visual understanding with advanced language processing, making it suitable for a variety of multimodal AI applications.
Key Capabilities
- Vision-Language Integration: Processes and interprets both image and text inputs, enabling comprehensive multimodal understanding.
- Large Context Window: Features a substantial context length of 32768 tokens, allowing for the handling of extensive visual and textual information.
- Instruction Following: Inherits strong instruction-following capabilities from its base model, facilitating precise responses to user prompts.
Good For
- Multimodal AI Applications: Ideal for tasks that require simultaneous processing of visual data and natural language.
- Image-to-Text Generation: Can be utilized for generating descriptions or answering questions based on image content.
- Interactive AI Systems: Suitable for developing conversational agents that can understand and respond to visual cues.