CountingSheep/vev-4b
CountingSheep/vev-4b is a 4.5 billion parameter multimodal decision model developed by CountingSheep, built upon Qwen/Qwen3.5-4B. It is designed to answer yes/no, multiple-choice, or graded questions about text, JSON records, or images, providing probability scores for each answer. This model excels at visual question answering and structured decision-making, serving as a Jev-like model with added image understanding capabilities, making it suitable for tasks like UI state analysis or real-time game interaction.
Loading preview...
vev-4b: Multimodal Decision Model with Image Understanding
vev-4b is a 4.5 billion parameter multimodal decision model developed by CountingSheep, extending the Qwen/Qwen3.5-4B architecture. It specializes in answering structured questions (yes/no, multiple-choice, graded) about diverse inputs including text, JSON records, and images, providing probability distributions for all possible answers in a single forward pass. This model replicates the request format of TypeSafe's Jev API but adds crucial visual processing capabilities, allowing it to interpret screenshots and other visual data.
Key Capabilities
- Multimodal Input Processing: Handles text, JSON, and images (screenshots) to answer questions.
- Probabilistic Decision Making: Provides probability scores for each answer option, enabling nuanced decision support.
- Jev-like API Compatibility: Uses the
/v1/systemonerequest format, familiar to users of TypeSafe's Jev. - Visual Question Answering: Achieves high accuracy on tasks like identifying UI states in app screenshots (e.g., 0.928 for switch/checkbox status, 0.951 for text input fields).
- Real-time Interaction: Demonstrated capability in scenarios like playing games by interpreting screenshots and making rapid decisions.
Good for
- Automated UI Testing: Analyzing screenshots to determine application states or identify errors.
- Content Moderation: Evaluating images against written safety policies.
- Structured Data Extraction: Answering specific questions from text or JSON records with probabilistic confidence.
- Interactive Agents: Developing agents that can perceive visual information and make structured decisions, such as in gaming or robotic control.
- Research in Multimodal AI: Exploring decision-making processes in models that integrate visual and textual understanding.