laion/exp-syh-r2egym-askllm-constrained_glm_4-7_traces_jupiter_cleaned-exp-gfi-swesmith-random-fi

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 22, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

This model is an 8 billion parameter language model, fine-tuned from Qwen/Qwen3-8B. It was trained on specific datasets related to 'exp-syh-r2egym-askllm-constrained_glm_4.7_traces_jupiter_cleaned' and 'exp-gfi-swesmith-random-filtered-10K_glm_4.7_traces_jupiter', suggesting a specialization in processing or generating content related to these trace datasets. With a context length of 32768 tokens, it is designed for tasks requiring extensive contextual understanding within its fine-tuned domain.

Loading preview...

Overview

This model is a fine-tuned version of the Qwen3-8B architecture, developed by laion. It has 8 billion parameters and supports a context length of 32768 tokens. The fine-tuning process utilized specific datasets: /e/data1/datasets/playground/ot/hf_hub/datasets--DCAgent--exp-syh-r2egym-askllm-constrained_glm_4.7_traces_jupiter_cleaned/snapshots/d13cd4ded646d8380dc70005a25fadeae9836514_thinking_preprocessed and /e/data1/datasets/playground/ot/hf_hub/datasets--DCAgent--exp-gfi-swesmith-random-filtered-10K_glm_4.7_traces_jupiter/snapshots/cb971aef68078d7bd025e0d8c33040bba180d914_thinking_preprocessed.

Training Details

The model was trained with a learning rate of 4e-05, a total training batch size of 96 (across 32 devices with 3 gradient accumulation steps), and for 7 epochs. It utilized the AdamW_TORCH_FUSED optimizer and a cosine learning rate scheduler with a 0.1 warmup ratio. The training was conducted using Transformers 4.57.6, Pytorch 2.9.1+cu130, Datasets 4.7.0, and Tokenizers 0.22.2.

Potential Use Cases

Given its fine-tuning on specific trace datasets, this model is likely best suited for applications that involve:

  • Processing or analyzing data similar to the exp-syh-r2egym-askllm-constrained_glm_4.7_traces_jupiter_cleaned dataset.
  • Tasks related to the exp-gfi-swesmith-random-filtered-10K_glm_4.7_traces_jupiter dataset, potentially in areas like log analysis, system tracing, or specific domain-related text generation/understanding.