namkoong-lab/LatentGym_Qwen3-8B_1episode_SingleLatent_number_guessing

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The namkoong-lab/LatentGym_Qwen3-8B_1episode_SingleLatent_number_guessing model is an 8 billion parameter Qwen3-based language model, fine-tuned using the GRPO algorithm. It is specifically optimized for single-latent training within the 'number_guessing' environment as part of the LatentGym testbed. This model is designed to excel at tasks related to number guessing, demonstrating specialized performance in this particular domain.

Loading preview...

Model Overview

This model, namkoong-lab/LatentGym_Qwen3-8B_1episode_SingleLatent_number_guessing, is an 8 billion parameter variant of the Qwen3 architecture, developed by namkoong-lab. It has been fine-tuned using the GRPO (Generalized Reinforcement Learning with Policy Optimization) algorithm, a method often employed for optimizing models in specific environments.

Key Capabilities & Training

This model is a specialized component of the LatentGym testbed, focusing on single-latent training within the number_guessing environment. During its training, the model was exposed to the set_of_3 latent configuration for the number guessing task. Key training hyperparameters include:

  • Base Model: Qwen/Qwen3-8B
  • Algorithm: GRPO
  • Learning Rate: 5e-07 with a constant_with_warmup schedule
  • Episodes per Trajectory (N): 1
  • Max Generation Length: 64 tokens
  • Sampling: Temperature (T)=0.8, top-p=0.95

Use Cases

This model is particularly suited for research and development in reinforcement learning environments, specifically for tasks involving number guessing where a single latent variable is critical. Its fine-tuned nature makes it a strong candidate for evaluating GRPO performance and understanding model behavior in constrained, latent-driven tasks.