kulia-moon/Lily-Qwen1.5-0.5B

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.6BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 10, 2025License:mitArchitecture:Transformer Open Weights Featherless Exclusive Warm

Lily-Qwen1.5-0.5B is a 0.6 billion parameter language model developed by kulia-moon, fine-tuned from the Qwen1.5-0.5B architecture. It features SwiGLU activation, attention QKV bias, and group query attention, supporting a 32K token context length. This model is optimized for natural language processing tasks, excelling in text generation, conversational dialogue, and language understanding, making it suitable for chatbots and content creation.

Loading preview...

Model Overview

Lily-Qwen1.5-0.5B is a fine-tuned version of the Qwen1.5-0.5B model, developed by kulia-moon. It leverages the Transformer architecture with SwiGLU activation, attention QKV bias, and group query attention. This model is designed to provide enhanced performance in various natural language processing tasks.

Key Capabilities

  • Text Generation: Capable of generating coherent and contextually relevant text based on prompts.
  • Conversational Dialogue: Optimized for interactive chat applications, acting as a friendly assistant.
  • Language Understanding: Excels in comprehending and processing natural language inputs.
  • Multilingual Support: Offers improved capabilities for handling multiple languages, including translation tasks.
  • Extended Context Length: Supports a stable context window of up to 32,768 tokens.

Ideal Use Cases

  • Chatbots and Virtual Assistants: Suitable for creating interactive and responsive conversational agents.
  • Content Creation: Can assist in generating various forms of text content, such as stories or articles.
  • Interactive Applications: Applicable in scenarios requiring dynamic language processing and understanding.

Limitations

  • Context Length: While extended, 32K tokens may still be insufficient for extremely long documents.
  • GQA Support: Lacks General Question Answering (GQA) support for most model sizes, which might affect performance on complex queries.
  • Common Sense: May occasionally struggle with nuanced human behavior or real-world context.