SaniaKhalid/tinyllama-peft-merged

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.1BQuant:BF16Context Size:2kPublished:Sep 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

SaniaKhalid/tinyllama-peft-merged is a 1.1 billion parameter TinyLlama model, fine-tuned using PEFT LoRA and fully merged for production readiness. This model, developed by Sania Khalid, is designed for causal language modeling tasks with a context length of 2048 tokens. It offers a streamlined inference experience as it requires no PEFT setup, making it easy to load and use directly.

Loading preview...

Overview

SaniaKhalid/tinyllama-peft-merged is a 1.1 billion parameter TinyLlama model that has been fine-tuned using PEFT (Parameter-Efficient Fine-Tuning) LoRA and subsequently merged. This merging process results in a production-ready model that does not require separate PEFT adapters for inference, simplifying deployment.

Key Capabilities

  • Simplified Inference: The model is fully merged, meaning users can load and run it directly without needing to manage PEFT adapters.
  • Causal Language Modeling: Designed for generating text based on a given prompt.
  • Compact Size: At 1.1 billion parameters, it offers a balance between performance and resource efficiency.
  • Standard Format: Provided in PyTorch Safetensors format with FP16 precision.
  • Context Window: Supports a context length of 2048 tokens.

Good For

  • Developers looking for a small, efficient language model for various text generation tasks.
  • Applications where ease of deployment and reduced inference complexity are crucial.
  • Educational or research purposes involving fine-tuned TinyLlama models without PEFT overhead.