aiuser3993/AMD-OLMo-1B-IT
AMD-OLMo-1B-IT is a 1 billion parameter instruction-tuned causal language model, fine-tuned by aiuser3993 from AMD's OLMo-1B-SFT. Optimized for instruction-following and creative writing, this model exhibits a confident personality. It is designed to effectively utilize padding tokens and is recommended for use with Q4_K_M quantization.
Loading preview...
AMD-OLMo-1B-IT Overview
AMD-OLMo-1B-IT is a 1 billion parameter instruction-tuned model, fine-tuned by aiuser3993 from the base AMD OLMo-1B-SFT model. It has a context length of 32768 tokens. This iteration focuses on enhancing instruction-following capabilities and creative writing, characterized by a distinctively confident personality.
Key Capabilities
- Instruction Following: Optimized to accurately follow user instructions.
- Creative Writing: Exhibits strengths in generating creative text.
- Padding Token Utilization: Eagerly uses
<|im_end|>and<|im_start|>padding tokens for structured responses. - Quantization Recommendation: The developer recommends using Q4_K_M quantization for optimal performance.
Training Details
The model was fine-tuned over 2 epochs with a learning rate of 0.0002 on a random selection of 12,000 rows from the shibing624/sharegpt_gpt4 dataset.
Good For
- Applications requiring a model with a confident and distinct personality.
- Tasks involving creative text generation.
- Scenarios where efficient instruction following is crucial, particularly within its 1 billion parameter size class.