laion/exp-syh-r2egym-askllm-constrained_glm_4-7_traces_jupiter_cleaned-exp-gfi-swesmith-random-fi
This model is an 8 billion parameter language model, fine-tuned from Qwen/Qwen3-8B. It was trained on specific datasets related to 'exp-syh-r2egym-askllm-constrained_glm_4.7_traces_jupiter_cleaned' and 'exp-gfi-swesmith-random-filtered-10K_glm_4.7_traces_jupiter', suggesting a specialization in processing or generating content related to these trace datasets. With a context length of 32768 tokens, it is designed for tasks requiring extensive contextual understanding within its fine-tuned domain.
Loading preview...
Overview
This model is a fine-tuned version of the Qwen3-8B architecture, developed by laion. It has 8 billion parameters and supports a context length of 32768 tokens. The fine-tuning process utilized specific datasets: /e/data1/datasets/playground/ot/hf_hub/datasets--DCAgent--exp-syh-r2egym-askllm-constrained_glm_4.7_traces_jupiter_cleaned/snapshots/d13cd4ded646d8380dc70005a25fadeae9836514_thinking_preprocessed and /e/data1/datasets/playground/ot/hf_hub/datasets--DCAgent--exp-gfi-swesmith-random-filtered-10K_glm_4.7_traces_jupiter/snapshots/cb971aef68078d7bd025e0d8c33040bba180d914_thinking_preprocessed.
Training Details
The model was trained with a learning rate of 4e-05, a total training batch size of 96 (across 32 devices with 3 gradient accumulation steps), and for 7 epochs. It utilized the AdamW_TORCH_FUSED optimizer and a cosine learning rate scheduler with a 0.1 warmup ratio. The training was conducted using Transformers 4.57.6, Pytorch 2.9.1+cu130, Datasets 4.7.0, and Tokenizers 0.22.2.
Potential Use Cases
Given its fine-tuning on specific trace datasets, this model is likely best suited for applications that involve:
- Processing or analyzing data similar to the
exp-syh-r2egym-askllm-constrained_glm_4.7_traces_jupiter_cleaneddataset. - Tasks related to the
exp-gfi-swesmith-random-filtered-10K_glm_4.7_traces_jupiterdataset, potentially in areas like log analysis, system tracing, or specific domain-related text generation/understanding.