FourOhFour/Vapor_7B
FourOhFour/Vapor_7B is a 7.6 billion parameter causal language model fine-tuned from Qwen/Qwen2.5-7B, designed for enhanced conversational capabilities across multiple languages. It leverages a diverse dataset including ShareGPT conversations and specialized instruction sets for reasoning and medical contexts. With a context length of 32768 tokens, it is optimized for complex dialogue and instruction-following tasks. The model integrates Liger plugins for improved performance and efficiency in its architecture.
Loading preview...
FourOhFour/Vapor_7B: Enhanced Conversational LLM
FourOhFour/Vapor_7B is a 7.6 billion parameter instruction-tuned language model built upon the Qwen/Qwen2.5-7B base. This model is designed to excel in conversational AI and instruction-following scenarios, supporting a wide array of languages including English, Chinese, French, Spanish, German, and more.
Key Capabilities & Training
Vapor_7B was fine-tuned on a curated collection of datasets, primarily focusing on ShareGPT-formatted conversations. Notable datasets include:
- PocketDoc/Dans-MemoryCore-CoreCurriculum-Small: Enhances general knowledge and conversational flow.
- NewEden/Kalo-Opus-Instruct-22k-Refusal-Murdered: Improves instruction adherence and refusal handling.
- Epiculous/Synthstruct-Gens-v1.1-Filtered-n-Cleaned: Refines synthetic data integration.
- Nitral-AI/Reasoning-1shot_ShareGPT: Boosts reasoning abilities.
- Nitral-AI/Medical_Instruct-ShareGPT: Provides specialized medical instruction following.
The model utilizes a chatml conversation template and supports a substantial context length of 32768 tokens, making it suitable for extended dialogues and complex prompts. Training incorporated advanced techniques such as flash_attention and Liger plugins (e.g., liger_rope, liger_rms_norm, liger_swiglu) for optimized performance and efficiency.
Ideal Use Cases
- Multilingual Chatbots: Its broad language support makes it suitable for global applications.
- Complex Instruction Following: Excels in scenarios requiring detailed and nuanced responses based on instructions.
- Reasoning Tasks: Benefits from dedicated reasoning datasets for improved logical processing.
- Specialized Domains: The inclusion of medical instruction data suggests potential for applications in healthcare information systems.