LIF1014/ptdbench-Llama-3.2-1B-Instruct
LIF1014/ptdbench-Llama-3.2-1B-Instruct is a 1.23 billion parameter instruction-tuned causal language model developed by Meta, part of the Llama 3.2 collection. Optimized for multilingual dialogue, it excels in agentic retrieval and summarization tasks across languages like English, German, and Spanish. This model features an optimized transformer architecture with Grouped-Query Attention and a 32768-token context length, making it suitable for assistant-like chat applications and deployment in constrained environments.
Loading preview...
Model Overview
LIF1014/ptdbench-Llama-3.2-1B-Instruct is a 1.23 billion parameter instruction-tuned model from Meta's Llama 3.2 family, designed for multilingual dialogue and agentic applications. It utilizes an optimized transformer architecture with Grouped-Query Attention (GQA) and has a context length of 32768 tokens. The model was pretrained on up to 9 trillion tokens, incorporating knowledge distillation from larger Llama 3.1 models, and fine-tuned using Supervised Fine-Tuning (SFT), Rejection Sampling (RS), and Direct Preference Optimization (DPO) for alignment with human preferences.
Key Capabilities
- Multilingual Support: Officially supports English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai, with training on a broader language set.
- Dialogue Optimization: Specifically tuned for assistant-like chat, agentic retrieval, and summarization tasks.
- Long Context Handling: Features a 32768-token context length, demonstrating strong performance on long-context benchmarks like Needle in Haystack.
- Efficiency: The 1B parameter size, combined with GQA, makes it suitable for deployment in constrained environments such as mobile devices.
Intended Use Cases
- Assistant-like Chat: Ideal for conversational AI applications.
- Agentic Systems: Well-suited for knowledge retrieval and summarization agents.
- Mobile AI: Designed for deployment in highly constrained environments like mobile devices.
- Query and Prompt Rewriting: Can be used to refine and optimize user inputs.