cuong1692001/Terminal-qwen-data-science-trajs-long

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 24, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The cuong1692001/Terminal-qwen-data-science-trajs-long is an 8 billion parameter Qwen-based language model, fine-tuned from cuong1692001/Terminal-terminal_traj on the qwen_data_science dataset. This model is designed for data science-related tasks, leveraging its specialized training to enhance performance in this domain. It features a substantial 32,768 token context length, making it suitable for processing extensive data science queries and code.

Loading preview...

Overview

This model, cuong1692001/Terminal-qwen-data-science-trajs-long, is an 8 billion parameter language model built upon the Qwen architecture. It is a fine-tuned iteration of cuong1692001/Terminal-terminal_traj, specifically adapted using the qwen_data_science dataset. The model is configured with a significant context length of 32,768 tokens, allowing it to handle complex and lengthy inputs relevant to data science.

Training Details

The model underwent training with the following key hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: 1 (train), 8 (eval)
  • Optimizer: AdamW with default betas and epsilon
  • LR Scheduler: Cosine
  • Epochs: 5.0

Intended Use Cases

While specific details are pending, the fine-tuning on a "qwen_data_science" dataset suggests its primary utility lies in tasks related to data science. This could include code generation for data analysis, natural language processing for scientific texts, or assisting with data interpretation and modeling. Its large context window further supports handling extensive datasets or complex problem descriptions within this domain.