SaniaKhalid/tinyllama-peft-merged
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.1BQuant:BF16Context Size:2kPublished:Sep 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
SaniaKhalid/tinyllama-peft-merged is a 1.1 billion parameter TinyLlama model, fine-tuned using PEFT LoRA and fully merged for production readiness. This model, developed by Sania Khalid, is designed for causal language modeling tasks with a context length of 2048 tokens. It offers a streamlined inference experience as it requires no PEFT setup, making it easy to load and use directly.
Loading preview...
Overview
SaniaKhalid/tinyllama-peft-merged is a 1.1 billion parameter TinyLlama model that has been fine-tuned using PEFT (Parameter-Efficient Fine-Tuning) LoRA and subsequently merged. This merging process results in a production-ready model that does not require separate PEFT adapters for inference, simplifying deployment.
Key Capabilities
- Simplified Inference: The model is fully merged, meaning users can load and run it directly without needing to manage PEFT adapters.
- Causal Language Modeling: Designed for generating text based on a given prompt.
- Compact Size: At 1.1 billion parameters, it offers a balance between performance and resource efficiency.
- Standard Format: Provided in PyTorch Safetensors format with FP16 precision.
- Context Window: Supports a context length of 2048 tokens.
Good For
- Developers looking for a small, efficient language model for various text generation tasks.
- Applications where ease of deployment and reduced inference complexity are crucial.
- Educational or research purposes involving fine-tuned TinyLlama models without PEFT overhead.