beyoru/seul-preview

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 9, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

seul-preview is a 9 billion parameter agentic language model developed by beyoru, designed for long-horizon reasoning and reliable business tool use. It is specifically trained to maintain context across extended workflows, interact safely with enterprise tools, and improve through reinforcement learning with verifiable outcomes. This model excels at multi-turn planning and verifiable execution, making it suitable for complex automation tasks.

Loading preview...

Model Overview

seul-preview is a 9 billion parameter agentic language model developed by beyoru, focusing on long-horizon reasoning and reliable business tool use rather than just benchmark performance. Its core design principles emphasize maintaining context over extended workflows, safe interaction with enterprise tools, and improvement through reinforcement learning with verifiable outcomes.

Key Capabilities & Design Goals

  • Long-horizon reasoning: Designed to handle complex, multi-step tasks.
  • Reliable enterprise tool use: Optimized for safe and effective interaction with business applications.
  • Stable multi-turn planning: Capable of consistent planning across multiple conversational turns or actions.
  • Verifiable execution: Focuses on ensuring that actions and outcomes can be verified.
  • Efficient reinforcement learning: Built to improve performance through reinforcement learning based on environmental feedback.

Performance Highlights

On a 60-task subset of AutomationBench, seul-preview (9B) achieves a 28% partial credit and 7% strict pass rate, outperforming Ornith-1.0-9B. This represents a significant improvement over its base model, with RLVR training raising partial credit by 6 points and more than doubling the strict pass rate. This demonstrates the effectiveness of its reinforcement learning approach in enhancing overall accuracy and closing the performance gap with larger models.

Limitations

  • The current training process was conducted on a single domain.
  • Due to GPU availability constraints, the model was trained for approximately half of the planned steps, suggesting potential for further improvement.