CompassioninMachineLearning/Qwen-3-8b-intermediate-epoch-3-CPT-10k-urban-density-dataset

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

The BrandonHowe/Qwen3-8b-urban-qwen-20260920-full-CPT-merged-epoch-3 is an 8 billion parameter Qwen3-based language model, developed by BrandonHowe. This model was trained on the `CompassioninMachineLearning/urban_12738_cleaned` dataset, focusing on specific document exposures over three epochs. It is provided as a merged BF16 model, ready for direct loading without an adapter, and is intended for applications requiring specialized knowledge from its training data.

Loading preview...

Model Overview

This model, BrandonHowe/Qwen3-8b-urban-qwen-20260920-full-CPT-merged-epoch-3, is an 8 billion parameter variant based on the Qwen3 architecture. It represents a merged BF16 model from epoch 3.0, step 1134, and is designed for direct use without requiring an adapter.

Training Details

The model was trained using the CompassioninMachineLearning/urban_12738_cleaned dataset. The training regimen involved 10,072 distinct documents, with 2,000 repeat exposures per epoch, and utilized 200 disjoint validation documents. The merging process was performed using Unsloth's native save_pretrained_merged(save_method="merged_16bit"), resulting in validated BF16 weights packaged into eight safetensors shards.

Key Characteristics

  • Architecture: Qwen3-based, 8 billion parameters.
  • Training Data: Specialized training on the CompassioninMachineLearning/urban_12738_cleaned dataset.
  • Format: Standalone merged BF16 model, ready for immediate deployment.
  • Context Length: Supports a context length of 32768 tokens.

Intended Use

This model is suitable for applications that can leverage its specific training on the urban_12738_cleaned dataset. Users should evaluate its performance for tasks related to the dataset's domain, as training did not specifically establish an improvement in compassion and requires separate evaluation.