laion/exp-psu-swesmith-1K_glm_4-7_traces_jupiter__0-98__Qwen3-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 19, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The laion/exp-psu-swesmith-1K_glm_4-7_traces_jupiter__0-98__Qwen3-8B is an 8 billion parameter language model, fine-tuned from the Qwen3-8B architecture. This model was specifically trained on the /e/data1/datasets/playground/ot/hf_hub/datasets--DCAgent--exp-psu-swesmith-1K_glm_4.7_traces_jupiter/snapshots/24c8342833108c3a15a23b64f37b83ff7e65efa4_thinking_preprocessed dataset. It is designed for tasks related to its specific fine-tuning data, offering specialized performance within that domain.

Loading preview...

Model Overview

This model, laion/exp-psu-swesmith-1K_glm_4-7_traces_jupiter__0-98__Qwen3-8B, is an 8 billion parameter language model derived from the Qwen3-8B architecture. It has undergone a specific fine-tuning process to adapt its capabilities to a particular dataset.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3-8B.
  • Fine-tuning Dataset: Trained on the /e/data1/datasets/playground/ot/hf_hub/datasets--DCAgent--exp-psu-swesmith-1K_glm_4.7_traces_jupiter/snapshots/24c8342833108c3a15a23b64f37b83ff7e65efa4_thinking_preprocessed dataset.
  • Training Hyperparameters: Utilized a learning rate of 4e-05, a total batch size of 96 (with gradient accumulation), and trained for 7 epochs using the AdamW optimizer with a cosine learning rate scheduler.

Intended Use Cases

Given its fine-tuning on a specific dataset, this model is best suited for applications and research that align with the characteristics and domain of the exp-psu-swesmith-1K_glm_4.7_traces_jupiter dataset. Developers should consider its specialized training for tasks requiring nuanced understanding or generation within that particular data distribution.