espressovi/BODHI-qwen-3-maze-8b-rlvr

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 4, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

espressovi/BODHI-qwen-3-maze-8b-rlvr is an 8 billion parameter Qwen-3 based language model, developed by espressovi. This model is specifically fine-tuned using Reinforcement Learning (RL) for maze-solving tasks, building upon the `espressovi/BODHI-qwen-3-maze-8b-distil` base. It is optimized for applications requiring navigation and problem-solving within maze environments, leveraging its 32768 token context length.

Loading preview...

Model Overview

The espressovi/BODHI-qwen-3-maze-8b-rlvr is an 8 billion parameter language model, part of the BODHI project, developed by espressovi. It is based on the Qwen-3 architecture and has been specifically trained for maze-solving tasks. This model represents an RL-trained iteration, starting from the espressovi/BODHI-qwen-3-maze-8b-distil base.

Key Capabilities

  • Maze Solving: The primary capability of this model is its proficiency in navigating and solving maze-like problems, a result of its specialized Reinforcement Learning (RL) training.
  • Qwen-3 Architecture: Leverages the robust Qwen-3 base model, providing a strong foundation for its specialized task.
  • RL-Trained: Benefits from Reinforcement Learning, which typically enhances performance on specific, goal-oriented tasks like maze navigation.

Good For

  • Research in RL for LLMs: Ideal for researchers exploring the application of Reinforcement Learning to large language models for specific problem domains.
  • Maze Navigation Tasks: Suitable for applications or simulations that involve solving complex mazes or pathfinding challenges.
  • Specialized Problem Solving: Can be a valuable asset for use cases requiring an LLM to perform well on highly structured, rule-based problem-solving scenarios.