radlab/pLLama3.2-1B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 17, 2024License:llama3.2Architecture:Transformer Featherless Exclusive Cold

The radlab/pLLama3.2-1B model is a 1 billion parameter language model from the Llama 3.2 family, specifically fine-tuned and DPO-processed for the Polish language by radlab. It excels in generating precise Polish text, having been trained on a proprietary dataset of 650,000 Polish instructions and an additional 100,000 examples for DPO focusing on linguistic correctness. This model is optimized for Polish-specific natural language processing tasks, offering enhanced communication accuracy compared to its base version.

Loading preview...

radlab/pLLama3.2-1B: Polish Language Model

This model is part of the radlab/pLLama3.2 collection, a series of Llama 3.2 architecture models specifically trained for the Polish language. The 1 billion parameter version (pLLama3.2-1B) has undergone both fine-tuning and a DPO (Direct Preference Optimization) process to enhance its ability to communicate precisely in Polish.

Key Capabilities

  • Polish Language Specialization: Significantly improved performance and precision in Polish compared to the base Meta-Llama-3.2 models.
  • Extensive Polish Training Data: Fine-tuned on a proprietary dataset of approximately 650,000 Polish instructions, semi-automatically generated from public datasets.
  • Linguistic Correctness via DPO: Further refined using a 100,000-example DPO dataset, teaching the model to select grammatically correct and well-written Polish texts over those with errors.
  • Optimized for Polish NLP: Designed for applications requiring accurate and nuanced understanding and generation of Polish text.

Training Details

The training involved a two-stage process:

  1. Post-training/Fine-tuning: 5 epochs on the 650k Polish instruction dataset.
  2. DPO Retraining: 15,000 steps on the 100k dataset focused on correct Polish writing.

Recommended Parameters

For optimal performance, radlab suggests using a temperature of 0.5 and a repetition_penalty of 1.2.