laion/glm-4_6-stackexchange-tezos-32ep-131k
The laion/glm-4_6-stackexchange-tezos-32ep-131k model is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B. It was specifically trained on the open-athena/glm-4.6-stackexchange-tezos-32ep-131k dataset, indicating a specialization towards content from StackExchange related to Tezos. With a context length of 32768 tokens, this model is optimized for processing and generating text within the domain of Tezos-related discussions and information found on StackExchange.
Loading preview...
Model Overview
This model, laion/glm-4_6-stackexchange-tezos-32ep-131k, is an 8 billion parameter language model derived from the Qwen/Qwen3-8B architecture. It has been fine-tuned on a specialized dataset, open-athena/glm-4.6-stackexchange-tezos-32ep-131k, which suggests a focus on content from StackExchange pertaining to the Tezos blockchain.
Key Characteristics
- Base Model: Qwen/Qwen3-8B
- Parameter Count: 8 billion parameters
- Context Length: 32768 tokens, enabling the processing of extensive inputs.
- Specialization: Fine-tuned on Tezos-related StackExchange data, indicating potential proficiency in understanding and generating text within this specific technical domain.
Training Details
The model was trained with the following hyperparameters:
- Learning Rate: 4e-05
- Batch Size: A total training batch size of 16 (1 per device with 8 devices and 2 gradient accumulation steps).
- Optimizer: ADAMW_TORCH_FUSED with betas=(0.9, 0.98) and epsilon=1e-08.
- Scheduler: Cosine learning rate scheduler with a 0.1 warmup ratio.
- Epochs: 7.0 epochs.
Intended Use Cases
Given its fine-tuning on Tezos StackExchange data, this model is likely best suited for applications requiring:
- Tezos-specific Q&A: Answering questions related to the Tezos blockchain, smart contracts, and ecosystem.
- Content Generation: Creating or summarizing technical discussions and documentation pertinent to Tezos.
- Information Retrieval: Extracting specific details from large bodies of Tezos-related text.
Further information regarding specific intended uses, limitations, and detailed training/evaluation data is noted as needing more documentation.