GestaltLabs/Ornstein3.8-27B
GestaltLabs/Ornstein3.8-27B is a 27 billion parameter vision-language fine-tune of Qwen/Qwen3.8-27B, built on the Qwen3_5ForConditionalGeneration architecture. This model integrates a SigLIP-style vision tower, enabling native image and video processing within its 262,144 token context window. It is specifically designed for multimodal tasks, demonstrating strong performance on benchmarks like GSM8K with an estimated 96.51% accuracy.
Loading preview...
Ornstein3.8-27B: A Multimodal Qwen3.8-27B Fine-tune
Ornstein3.8-27B is a 27 billion parameter vision-language model developed by GestaltLabs, fine-tuned from the Qwen/Qwen3.8-27B base model. It leverages the Qwen3_5ForConditionalGeneration architecture, which features interleaved linear and full attention (Gated DeltaNet) and supports a substantial context length of 262,144 tokens. A key differentiator is its integrated SigLIP-style vision tower, allowing for native processing of image and video inputs.
Key Capabilities
- Multimodal Understanding: Processes both text and visual inputs (images and video) natively.
- Extended Context: Supports a large context window of 262,144 tokens, beneficial for complex, long-form interactions.
- Strong Reasoning: Achieves an estimated 96.51% accuracy on the GSM8K benchmark, indicating robust mathematical and reasoning capabilities.
- Qwen3.8 Base: Inherits the strong language understanding and generation abilities of the Qwen3.8-27B model.
Good For
- Vision-Language Tasks: Ideal for applications requiring the interpretation of images or video alongside text.
- Complex Problem Solving: Suitable for tasks demanding strong reasoning, such as mathematical word problems.
- Research and Development: An early merge checkpoint for exploring advanced multimodal AI, with planned future enhancements using RL environments and energy-based fine-tuning.
This model is released under the Apache 2.0 license, inherited from its Qwen 3.8 base.