laion/exp-gfi-swesmith-random-filtered-10K_glm_4_7_traces_jupiter_cleaned

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

laion/exp-gfi-swesmith-random-filtered-10K_glm_4_7_traces_jupiter_cleaned is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B. This model was trained on the /data/cat/ws/befe330h-befe330h-otagent/huggingface/hub/datasets--DCAgent--exp-gfi-swesmith-random-filtered-10K_glm_4.7_traces_jupiter_cleaned/snapshots/35ab4be85fb10668d8de78ade5d36e0e9f546f6e_thinking_preprocessed dataset, suggesting a specialization in processing or generating content related to specific data traces. With a context length of 32768 tokens, it is designed for tasks requiring extensive contextual understanding.

Loading preview...

Model Overview

This model, exp-gfi-swesmith-random-filtered-10K_glm_4_7_traces_jupiter_cleaned, is an 8 billion parameter language model based on the Qwen3-8B architecture. It has been fine-tuned on a specific dataset: /data/cat/ws/befe330h-befe330h-otagent/huggingface/hub/datasets--DCAgent--exp-gfi-swesmith-random-filtered-10K_glm_4.7_traces_jupiter_cleaned/snapshots/35ab4be85fb10668d8de78ade5d36e0e9f546f6e_thinking_preprocessed.

Training Details

The fine-tuning process involved the following key hyperparameters:

  • Learning Rate: 4e-05
  • Batch Size: A total training batch size of 16 (with train_batch_size: 1 and gradient_accumulation_steps: 2 across 8 GPUs).
  • Optimizer: ADAMW_TORCH_FUSED with betas=(0.9, 0.98) and epsilon=1e-08.
  • Scheduler: Cosine learning rate scheduler with a 0.1 warmup ratio.
  • Epochs: Trained for 7.0 epochs.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3-8B.
  • Parameter Count: 8 billion parameters.
  • Context Length: Supports a context window of 32768 tokens.
  • Specialization: The specific training dataset suggests a focus on tasks related to processing or analyzing "thinking preprocessed" data traces, potentially for agent-based systems or complex reasoning tasks.

Potential Use Cases

Given its fine-tuning on a specialized dataset, this model could be particularly useful for:

  • Analyzing and generating text based on specific data trace patterns.
  • Applications requiring deep contextual understanding from structured or semi-structured trace data.
  • Tasks within domains where the original training dataset's characteristics are relevant.