shirley09221/qwen3-8b-full-pretrain-junk-tweet-1m-en

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The shirley09221/qwen3-8b-full-pretrain-junk-tweet-1m-en model is an 8 billion parameter language model based on the Qwen3-8B architecture. It has been fine-tuned specifically on the junk_tweet_1m_en dataset, indicating a specialization in processing or generating content related to 'junk tweets' in English. This model is distinct due to its targeted fine-tuning on a specific, potentially noisy, social media dataset, which could make it suitable for tasks involving informal or low-quality text analysis.

Loading preview...

Model Overview

This model, shirley09221/qwen3-8b-full-pretrain-junk-tweet-1m-en, is an 8 billion parameter language model derived from the Qwen/Qwen3-8B base architecture. Its primary distinction lies in its fine-tuning process, which utilized the junk_tweet_1m_en dataset. This specialized training suggests an adaptation for understanding or generating content characteristic of informal, potentially low-quality, or 'junk' tweets in English.

Training Details

The model was trained with the following key hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: 1 (train), 8 (eval)
  • Optimizer: AdamW_Torch with default betas and epsilon
  • Scheduler: Cosine learning rate scheduler with 0.1 warmup steps
  • Epochs: 3.0
  • Devices: Trained across 8 GPUs

Potential Use Cases

Given its specific fine-tuning on a 'junk tweet' dataset, this model could be particularly relevant for:

  • Social Media Analysis: Tasks involving the classification, filtering, or understanding of informal and potentially low-quality content on platforms like Twitter.
  • Noise Robustness: Research or applications requiring a model that can handle noisy, ungrammatical, or colloquial text effectively.
  • Content Moderation: Identifying or categorizing specific types of undesirable content within social media streams.

Further details on intended uses, limitations, and comprehensive evaluation data are noted as needing more information in the original model card.