laion/a3-rl-laion_nemotron-gym-knowledge-web-search-mcqa-25-8B

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 4, 2026Architecture:Transformer Featherless Exclusive Cold

The laion/a3-rl-laion_nemotron-gym-knowledge-web-search-mcqa-25-8B model is an 8 billion parameter language model, based on the Qwen3-8B architecture, that has been fine-tuned using reinforcement learning (RL). It is specifically optimized for tasks involving knowledge retrieval, web search, and multiple-choice question answering (MCQA). This model leverages a fully-async RL training approach to enhance its performance in these specific domains, building upon a strong base SFT model.

Loading preview...

Model Overview

This model, laion/a3-rl-laion_nemotron-gym-knowledge-web-search-mcqa-25-8B, is an 8 billion parameter language model built on the Qwen3-8B architecture. It has undergone reinforcement learning (RL) fine-tuning using a fully-asynchronous, fan-out K=4 approach, specifically targeting improved performance in knowledge-intensive tasks.

Key Capabilities & Training

  • Reinforcement Learning (RL) Fine-tuning: The model was trained using SkyRL's fully-async method over 80 steps, with this specific checkpoint selected at step 25 based on EMA-best reward/avg_raw_reward.
  • Optimized for Specific Tasks: It is designed for applications requiring knowledge retrieval, web search, and multiple-choice question answering (MCQA), indicating a focus on factual accuracy and information synthesis.
  • Base Model: The RL training commenced from laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink, a strong supervised fine-tuned (SFT) base.
  • Training Traces Available: The training-time Daytona/Harbor rollouts are available as a companion dataset, open-athena/a3-rl-laion_nemotron-gym-knowledge-web-search-mcqa, providing transparency into its training process.

When to Use This Model

This model is particularly well-suited for use cases where robust performance in knowledge-based question answering, information extraction from web content, and accurate selection of answers from multiple choices is critical. Its RL-tuned nature suggests enhanced reasoning and factual grounding for these specific applications.