Yugong09/GeoGuess

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Apr 19, 2026Architecture:Transformer Featherless Exclusive Cold

Yugong09/GeoGuess is a 4 billion parameter vision-language model based on the Qwen3-VL-4B-Instruct architecture, developed by Yugong09. This model is designed for multimodal tasks, leveraging its vision capabilities to process and understand image inputs. With a context length of 32768 tokens, it excels in applications requiring both visual comprehension and natural language interaction.

Loading preview...

Yugong09/GeoGuess: A Vision-Language Model

Yugong09/GeoGuess is a 4 billion parameter vision-language model built upon the Qwen3-VL-4B-Instruct architecture. This model integrates robust visual understanding with advanced language processing, making it suitable for a variety of multimodal AI applications.

Key Capabilities

  • Vision-Language Integration: Processes and interprets both image and text inputs, enabling comprehensive multimodal understanding.
  • Large Context Window: Features a substantial context length of 32768 tokens, allowing for the handling of extensive visual and textual information.
  • Instruction Following: Inherits strong instruction-following capabilities from its base model, facilitating precise responses to user prompts.

Good For

  • Multimodal AI Applications: Ideal for tasks that require simultaneous processing of visual data and natural language.
  • Image-to-Text Generation: Can be utilized for generating descriptions or answering questions based on image content.
  • Interactive AI Systems: Suitable for developing conversational agents that can understand and respond to visual cues.