WideSeek-R1/WideSeek-SFT-4B
WideSeek-SFT-4B is a 4 billion parameter language model developed by WideSeek-R1, based on the Qwen3 architecture. It is supervised fine-tuned on agent-level multi-turn trajectories from the WideSeek-R1 SFT Data. This model specializes in handling both width-only and depth-only tasks, making it suitable for complex conversational AI and agentic workflows. With a context length of 32768 tokens, it can process extensive multi-turn interactions.
Loading preview...
WideSeek-SFT-4B Overview
WideSeek-SFT-4B is a 4 billion parameter language model built upon the Qwen3 architecture. Developed by WideSeek-R1, this model has undergone supervised fine-tuning using the proprietary WideSeek-R1 SFT Data. This dataset is specifically designed with agent-level multi-turn trajectories, focusing on both 'width-only' and 'depth-only' task types.
Key Capabilities
- Agentic Task Handling: Optimized for processing complex, multi-turn interactions typical in agent-based systems.
- Specialized Fine-tuning: Benefits from a unique dataset that targets distinct conversational patterns:
- Width-only tasks: Likely involves exploring multiple options or branches at a single decision point.
- Depth-only tasks: Suggests a focus on sequential reasoning and detailed exploration within a single path.
- Qwen3 Foundation: Leverages the robust capabilities of the Qwen3 base model.
- Extended Context: Supports a context length of 32768 tokens, enabling it to maintain coherence over long dialogues and complex task sequences.
Good For
- Developing AI agents that require nuanced multi-turn conversational abilities.
- Applications needing to manage tasks that involve either broad exploration of options or deep, sequential reasoning.
- Research into agentic AI behaviors and complex task execution.